GTFS Scorecard: plain-language quality for transit feeds

Role: Independent product designer, engineer, and operator · 2026

I run GTFS Scorecard, a live service that reads more than 2,100 published GTFS feed records and turns each one into a plain-language scorecard with a first recommended fix. Structural correctness comes from MobilityData's canonical validator; the scorecard adds freshness, rider-information, and realtime checks on top. Feed records are not necessarily distinct agencies, and no agency is known to have adopted a scorecard into its workflow.

Evidence boundary

Built
A live scorecard service with a read API, a GitHub Action, and a read-only MCP server
Observed
2,100+ curated feed records across 40+ countries, rechecked on a daily schedule
Not proved
Agency adoption, rider outcomes, or that any score changed a published feed

Selected measures

curated feed records in the registry; feed records are not necessarily distinct agencies
2,100+
countries with feed records in the registry
40+
scheduled recheck of every published feed record
Daily
agencies known to have adopted a scorecard into their workflow
0

GTFS Scorecard (opens in a new tab) reads a published GTFS feed and writes down what a rider would want fixed first. It is a live beta: feed records are rechecked on a daily schedule, every scorecard is a public page, and anyone can check a GTFS ZIP before publishing it. A scorecard describes a published file on a given day. It does not describe an agency's operations, and it does not claim the agency has read it.

Correctness is delegated

The scorecard does not re-litigate the specification. Structural correctness comes from MobilityData's canonical validator (opens in a new tab), the same tool feed publishers already run. What the scorecard adds is the layer above it: whether the feed is fresh, whether it carries the information a rider actually uses, and whether the realtime feed works. A feed can follow the specification (opens in a new tab) and still strand a rider.

What a grade will not do

  • Realtime is scored only when a usable realtime feed is configured. A missing realtime feed does not lower the grade.
  • The registry counts feed records, and the site says so. Feed records are not necessarily distinct agencies.
  • The first recommended fix is chosen for riders, not for specification completeness.
GTFS Scorecard home page with agency search, registry coverage counts, and the workflow steps.
The live home page: search an agency, open its latest scorecard, or check a GTFS ZIP before publishing it. · gtfsscorecard.org · live-service screenshot · captured August 17, 2026

A finding has to travel

A quality finding is only useful once it reaches the person who can fix the feed, and that person works somewhere else. So the same data ships four ways: the public scorecard pages, a read API, a GitHub Action that checks a feed in CI before it is published, and a read-only MCP server. Nobody has to ask my permission to carry a finding into their own workflow.

A published scorecard page for Unitrans, showing the grade, its coverage badges, and the note that it is not a compliance determination.
A published scorecard for Unitrans (ASUCD / City of Davis): the grade, what it covers, and, in its own footer, that it is not an official compliance determination. · gtfsscorecard.org · live-service screenshot · captured August 17, 2026

Paying for it in the open

Running the service costs single-digit dollars a month today, and the support page (opens in a new tab) says exactly that, along with what sponsorship would fund and the fact that there are no sponsors yet. The support model is published the same way the scores are: current state first, stated plainly, never rounded up.

What running it taught me

The code is public at github.com/ChelseaKR/gtfs-scorecard (opens in a new tab), and I wrote about what operating it changed in how I read public data in What GTFS Scorecard taught me about public data.

Outside review, and what it changed

The most useful review the project has had came as an argument to shut part of it down (opens in a new tab). A longtime open transit-data maintainer said the scoring belonged inside MobilityData’s canonical validator rather than in one more dashboard. He was substantially right. I named in the thread the tools I had duplicated, declined to push subjective letter grades into an official project where they would read as guidance, and narrowed this one to the handoff nobody else covers.

The person who produces the MRC de Joliette feed pushed back (opens in a new tab) on a recommendation to populate trip headsigns on loop routes. They were right, and the rule now credits that case instead of flagging it.

Upstream, reading the specifications closely enough to write rules produced a specification example fix (opens in a new tab) and a conformance-language clarification (opens in a new tab) merged into the Transit Operational Data Standard, and an awesome-transit listing (opens in a new tab). A second listing (opens in a new tab) and a Transitland feed-archival change (opens in a new tab) are open. Five small contributions, three of them merged, as of September 2026. That is the whole claim.

Use the live GTFS Scorecard

  • Public transit
  • GTFS
  • Civic data
  • Live service
  • The score was the easy part

    I built GTFS Scorecard to turn transit-data findings into repairs. It taught me that civic technology earns trust through explicit judgment, honest unknowns, preserved evidence, and proof that a change reached the source.