GTFS Scorecard: plain-language quality for transit feeds
Role: Independent product designer, engineer, and operator · 2026
I run GTFS Scorecard, a live service that reads more than 2,100 published GTFS feed records and turns each one into a plain-language scorecard with a first recommended fix. Structural correctness comes from MobilityData's canonical validator; the scorecard adds freshness, rider-information, and realtime checks on top. Feed records are not necessarily distinct agencies, and no agency is known to have adopted a scorecard into its workflow.
Evidence boundary
- Built
- A live scorecard service with a read API, a GitHub Action, and a read-only MCP server
- Observed
- 2,100+ curated feed records across 40+ countries, rechecked on a daily schedule
- Not proved
- Agency adoption, rider outcomes, or that any score changed a published feed
Selected measures
- curated feed records in the registry; feed records are not necessarily distinct agencies
- 2,100+
- countries with feed records in the registry
- 40+
- scheduled recheck of every published feed record
- Daily
- agencies known to have adopted a scorecard into their workflow
- 0
GTFS Scorecard (opens in a new tab) reads a published GTFS feed and writes down what a rider would want fixed first. It is a live beta: feed records are rechecked on a daily schedule, every scorecard is a public page, and anyone can check a GTFS ZIP before publishing it. A scorecard describes a published file on a given day. It does not describe an agency's operations, and it does not claim the agency has read it.
Correctness is delegated
The scorecard does not re-litigate the specification. Structural correctness comes from MobilityData's canonical validator (opens in a new tab), the same tool feed publishers already run. What the scorecard adds is the layer above it: whether the feed is fresh, whether it carries the information a rider actually uses, and whether the realtime feed works. A feed can follow the specification (opens in a new tab) and still strand a rider.
What a grade will not do
- Realtime is scored only when a usable realtime feed is configured. A missing realtime feed does not lower the grade.
- The registry counts feed records, and the site says so. Feed records are not necessarily distinct agencies.
- The first recommended fix is chosen for riders, not for specification completeness.

A finding has to travel
A quality finding is only useful once it reaches the person who can fix the feed, and that person works somewhere else. So the same data ships four ways: the public scorecard pages, a read API, a GitHub Action that checks a feed in CI before it is published, and a read-only MCP server. Nobody has to ask my permission to carry a finding into their own workflow.

Paying for it in the open
Running the service costs single-digit dollars a month today, and the support page (opens in a new tab) says exactly that, along with what sponsorship would fund and the fact that there are no sponsors yet. The support model is published the same way the scores are: current state first, stated plainly, never rounded up.
What running it taught me
The code is public at github.com/ChelseaKR/gtfs-scorecard (opens in a new tab), and I wrote about what operating it changed in how I read public data in What GTFS Scorecard taught me about public data.
Outside review, and what it changed
The most useful review the project has had came as an argument to shut part of it down (opens in a new tab). A longtime open transit-data maintainer said the scoring belonged inside MobilityData’s canonical validator rather than in one more dashboard. He was substantially right. I named in the thread the tools I had duplicated, declined to push subjective letter grades into an official project where they would read as guidance, and narrowed this one to the handoff nobody else covers.
The person who produces the MRC de Joliette feed pushed back (opens in a new tab) on a recommendation to populate trip headsigns on loop routes. They were right, and the rule now credits that case instead of flagging it.
Upstream, reading the specifications closely enough to write rules produced a specification example fix (opens in a new tab) and a conformance-language clarification (opens in a new tab) merged into the Transit Operational Data Standard, and an awesome-transit listing (opens in a new tab). A second listing (opens in a new tab) and a Transitland feed-archival change (opens in a new tab) are open. Five small contributions, three of them merged, as of September 2026. That is the whole claim.
Writing about this work

The score was the easy part
I built GTFS Scorecard to turn transit-data findings into repairs. It taught me that civic technology earns trust through explicit judgment, honest unknowns, preserved evidence, and proof that a change reached the source.