How BePünktli measures punctuality
Every number on this site comes from two public datasets that were never designed to be joined to each other. This page explains what they contain, how we connect them, how often that connection fails, and what we therefore do not claim. It is written for people who did not build the system.
1. What a measurement is
A vehicle's journey is a sequence of stops. At each stop the timetable promises an arrival time and a departure time, and the operator afterwards reports what actually happened.
A delay is actual − scheduled, in seconds, signed. A train three minutes late is +180. A bus
two minutes early is −120, and it stays negative: we never round an early departure up to "on
time", because departing early is its own kind of failure — it means passengers who arrived on time
for the published departure missed it.
Punctuality is then the share of arrivals that were less than three minutes late. Three minutes is the Swiss industry convention, and we publish the five- and fifteen-minute figures beside it so the choice of threshold is visible rather than load-bearing.
2. The two sources
istdaten — what happened
The Swiss public-transport data portal publishes a daily file of actual stop events (Ist-Daten): one row per stop of every journey operated in the country, roughly two million rows a day. We hold every day from 1 January 2022 onward. Up to 21 September 2026 that is 3,743,695,749 rows covering 226,428,936 journeys over 1,721 operating days.
That is four days short of the calendar's 1,725, and eight days are damaged in total: four were never published at all — those are the four missing from the count above — and four more were published truncated, so they are present but incomplete. All eight are permanent; no backfill exists anywhere. We count them as damaged rather than as days on which nothing ran, and any period containing one says so on the page.
GTFS — what was promised
The same publisher issues the national timetable in the GTFS format, roughly twice a week. Each publication is a complete snapshot of the timetable as it stood that day, including changes that were only decided a few days earlier — engineering closures, replacement buses, and seasonal variations.
That cadence matters. A journey has to be compared against the timetable that was in force on the day it ran, not against today's. Using a timetable four months stale costs about five percentage points of matching, which we measured directly: a period that ran on one old snapshot matched 95.5% of its journeys where properly-dated neighbours matched 98.6%.
We therefore keep the snapshots as an archive — 214 of them, back to 2022 and forward into next year's timetable — and every operating day is matched against the snapshot that was current on that date.
3. The columns we use
istdaten's columns are German. Here is what each one means, what we do with it, and where it is incomplete.
| column | meaning | how we use it | where it is missing or awkward |
|---|---|---|---|
BETRIEBSTAG | operating day | the unit everything is grouped by | a journey that runs past midnight belongs to the day it started |
FAHRT_BEZEICHNER | journey reference, the operator's id for this run | the first thing we try to match on | reused across days, so it is only an identifier together with the date. 20 operators publish a different one in the timetable than their vehicles report |
BETREIBER_ID | operator, e.g. 85:11 (SBB) | identifies the company | the 85: prefix is Switzerland; Austrian operators begin 81:, which is why we never strip a fixed prefix |
LINIEN_ID | line identifier | identifies the line | means two different things. For bus and tram it is a self-describing code like 85:801:701. For rail it is a train number, which is reassigned between timetable years |
LINIEN_TEXT | the label a passenger sees | display, and a weak matching signal | labels are not unique: five different lines are called "231" |
PRODUKT_ID | mode — train, bus, tram, … | grouping and product scope | sometimes empty, and not only for obscure services — ordinary regional trains appear with no mode. An empty mode is treated as mainline rail and kept, never filtered out |
BPUIC | stop code | identifies the stop | two formats. Seven digits is a station; nine characters is that station plus a two-character platform code, which need not be numeric. Operators migrated during 2026, so ~80% of recent rows carry the longer form against 10.9% of the archive |
HALTESTELLEN_NAME | stop name as the operator writes it | never used for identity | several operators publish none at all, and one operator went from 1.2% filled to 100% filled in the course of a single month. Stop names come from a registry instead |
ANKUNFTSZEIT / ABFAHRTSZEIT | scheduled arrival / departure | the promise | absent at a journey's first stop (no arrival) and last stop (no departure) |
AN_PROGNOSE / AB_PROGNOSE | reported arrival / departure | the outcome | present only when the status below says so |
AN_PROGNOSE_STATUS / AB_PROGNOSE_STATUS | how the reported time was obtained | decides whether the row is a measurement at all | see below |
FAELLT_AUS_TF | cancelled | a failure, counted in the denominator | — |
DURCHFAHRT_TF | passed through without stopping | excluded — it is not a stop | — |
ZUSATZFAHRT_TF | an extra, unscheduled run | excluded from rates — there is no promise to measure it against | it is matched to the timetable where possible, because a fifth of them turn out to be in it |
The status column is the one that matters most
AN_PROGNOSE_STATUS takes one of four values, and they are not four grades of the same thing:
REAL— an arrival event fired at the stop and the system sent the time it happened.PROGNOSE— no event fired; the time is the last forecast the system held, which is what the departure board showed. Counted, separately — see §6.UNBEKANNT— the system has no deviation for this stop at all.- empty — nothing was reported.
This column is the single easiest thing to get wrong, and getting it wrong inverts a ranking. A
natural-looking test — "is the status not REAL?" — is true for UNBEKANNT and for empty, so
writing it that way quietly scores every stop nobody reported on as punctual, and the operators with
the worst telemetry come out the most reliable. We never score a stop with no time on it; it is
excluded from the rate and counted in the coverage figure instead.
Those statuses decide what each stop event becomes. For a recent representative month, the national outcome of every stop event we were entitled to measure was:
| events | share | scored? | |
|---|---|---|---|
an arrival event fired (REAL) | 62,329,333 | 93.53% | yes |
a forecast was held (PROGNOSE) | 1,459,133 | 2.19% | yes |
| the forecast was the timetable | 556,015 | 0.83% | no — §6 |
| cancelled | 528,483 | 0.79% | yes, always as late |
| no time at all | 1,761,108 | 2.64% | no |
| implausible value | 6,320 | 0.01% | no |
The third row is the rule §6 sets out at length: a forecast filed at the scheduled second, in a journey carrying at least one other, is the plan copied rather than a measurement, so it is in no rate's numerator and in no rate's denominator.
So a headline punctuality figure for that month is computed over 96.51% of the stop events it could have used — and 97.71% of what it counted fired an arrival event. We print both.
4. The hard part: which timetable entry is this journey?
The two datasets describe the same vehicle movements and, for most of our history, share no reliable identifier.
The timetable's original_trip_id field — which names the journey reference each trip corresponds
to — only appears in feeds published from 14 December 2024 onward. For everything before that,
roughly three years of the archive, the timetable simply does not say which journey its trips are.
And even where the field exists it does not solve the problem: in a recent month it matched
1,971,546 journeys while 2,052,286 had to be matched another way, because those operators
publish a different reference in the timetable than their vehicles report.
So we match on structure: the shape of the journey itself.
A journey's stop pattern — which stops, in which order, at which times — is close to a fingerprint. Two different trains rarely call at the same stops at the same times on the same day. We use that, in five steps, taking the first that gives exactly one answer:
| step | what it compares | how reliable |
|---|---|---|
| 1. Published reference | the journey reference and the date, refusing if more than one trip matches | 98.5% correct on operators with stable references |
| 2. Exact pattern | the full multiset of stops and times, within the same day and operator | 0 errors in 416,020 decisions |
| 3. Contained pattern | one journey's stops contained in the other's, at least 3 stops and at least 70% overlap | 0 errors in 7,647 decisions |
| 4. Ends and first departure | first stop, last stop, first departure time | 9 errors in 23,210 decisions — 99.96% |
| 5. Journey number | the operator's own run number, then checked against the stop pattern | used where a legacy identifier is all there is |
Two rules make this trustworthy rather than merely clever.
If more than one timetable entry fits, we record no match. Choosing the most likely candidate would be the single largest source of silent error in a system like this, and a wrong match is worse than no match: it scores one line's delay against another line's schedule, and every figure downstream inherits it. Journeys that fit two entries are marked ambiguous and excluded.
Every match records which step produced it. A number is only as good as the weakest step behind it, so the step is stored with the journey and can be audited afterwards.
How often this works
96.50% of all 225,148,610 journeys are matched to their timetable entry. By year:
| 2022 | 2023 | 2024 | 2025 | 2026 |
|---|---|---|---|---|
| 94.9% | 95.7% | 96.1% | 97.4% | 98.8% |
The older years are lower because they have no published reference at all and rely entirely on structure. The unmatched remainder is not a random sample — it is weighted towards journeys whose timetable had changed, which means towards disrupted service. That is why we publish the matching rate next to every figure rather than only the figure.
5. The other hard part: which line is this?
Knowing which journey a row belongs to is not the same as knowing which line. The two datasets name lines in two incompatible ways, and which one applies depends on the mode.
- Bus and tram — together 92.8% of all stop events — carry a self-describing line code like
85:801:701. That is a perfectly good identifier, and it is not a timetable route id. - Rail carries a train number, which identifies one run and not a line — and which is reassigned between timetable years, so the same number can mean different lines in different years.
The timetable, meanwhile, names everything by route id such as 96-701-j23-1, where the j23
part is the timetable year.
This has a consequence worth stating plainly, because it looks like a data gap and is not one: an inventory that compares the two sides directly finds thousands of timetable routes with "no observations at all", and most of that is the mismatch rather than missing service. PostAuto's line 101 appears in our data twice — as an observed line with 1,538 runs a week, and as a route with no observations and 1,629 scheduled trips a week. It is one healthy line running at 94% of its schedule.
We resolve this by deriving the connection rather than assuming it: the journeys matched in section 4 tell us which timetable route a line's runs actually belong to. That derived connection is only accepted when five conditions hold, and it is refused otherwise:
- the line's journeys reach exactly one route in that timetable year;
- that route is reached by exactly one line, so no route's scheduled service is counted twice;
- there are enough journeys behind the pairing to be more than a coincidence;
- at most a tenth of the line's year ran somewhere the pairing does not account for, and that shortfall is recorded;
- the stops the line was observed serving and the stops the route is scheduled to serve actually agree.
The fifth condition is what separates associated from identical. Three quarters of accepted pairings match perfectly — every observed stop is a scheduled stop and vice versa. The ones we reject include replacement-service routes that shared a single stop with a line, and routes that turned out to be a fragment of a longer line, which would have made that line's schedule look smaller than it is.
0.68% of stop events cannot be attributed to a named line. Those rows keep everything else — station, times, delay — and count in national and per-operator figures; they simply do not appear on a line's page.
6. Coverage, and why there are three numbers
A punctuality rate is only meaningful next to the share of service it was computed over. We publish three, and the figure we show is always the worst of them, never the average:
| axis | question |
|---|---|
| days | did the publisher release the day at all? |
| journeys | what share of the day's journeys reached a timetable entry? |
| measurement | what share of the stop events we were entitled to measure carry a real reading? |
They are independent, and any of the three can be the binding one. A recent month had all its days, matched 98.6% of its journeys, and still measured only 94.3% of its stop events — so it publishes at 94.3%. Averaging the three would have shown 97.6% and hidden the axis that was actually thin.
Where coverage is below the threshold we consider publishable, we show the state rather than a greyed-out number. A figure computed over a tenth of a line's service is not a cautious estimate; it is a different claim, and the honest thing is not to make it.
Two grades of time, both counted, and the difference always visible
Swiss operators file a row for every stop event, and each row carries a status saying what kind of time it holds. The Swiss realisation spec for that field defines two that matter:
REAL— an arrival event fired at the stop: a door-opening signal, an entry or exit loop, and the operator's system sent the time it happened.PROGNOSE— no such event fired. The time is the last forecast the system held for that stop, which is also what was on the departure board.
We count both, and we keep them apart. The reason is that the missing event is a property of the
equipment at a stop, not of the journey: a stop with no door signal and no loop stays PROGNOSE
however plainly the vehicle served it. Until September 2025, Geneva's Transports Publics Genevois
filed 5 million stop events a month that way and not one with an event — while running one of the
country's largest networks. Scoring only the confirmed grade would have measured which operators own
which hardware.
We checked before deciding. Those forecasts are not the timetable restated: 99.5% of them differ from the planned time, and when TPG and Lausanne's tl each began detecting arrival events — on different dates, two years apart — the punctuality distribution did not step. Their systems were tracking vehicles well before they could confirm them.
What we do not do is hide the difference. Every figure can say what share of it fired an event, because two things are only visible to someone who can see that split: a vehicle may have driven past a stop without the system noticing, and where no forecast exists at all the value sent is the timetable — which we never score, because a timetable is not a measurement of anything.
That last one is a rule with a shape, and it is worth stating exactly, because it is the difference between a punctuality figure and a copy of the plan. The standard says that when no forecast is available the value to send is the scheduled time itself. A stop reported that way is not evidence that anything arrived on time; it is the timetable, and scored against the timetable it is perfect by construction. So a forecast equal to its scheduled time to the second is not counted as a measurement — it leaves the rate and joins the coverage figure, exactly as a stop with no time at all does.
With one exception, because a vehicle really can arrive at the scheduled second. We only apply the rule where the same journey has at least two such stops. That is not a judgement call: arrivals that did fire an event and landed exactly on the scheduled second — which are genuine by definition — are the only such stop in their journey 43–67% of the time, while forecast stops filed at the scheduled second are isolated only 8–14% of the time and typically come in runs covering a whole journey. A single coincidence is kept. A journey reported entirely as its own timetable is not.
Every figure says how much of it this removed, and the effect is small nationally and large for the few feeds it is about: a tenth of a point of punctuality and under a point of coverage across the country, against whole operator-months that stop being publishable. One city bus operator ran seven consecutive months in which essentially every arrival it filed was the scheduled second exactly, and published 100.00% punctuality at 100% coverage for each of them. It now publishes nothing for those months, which is the honest answer: we cannot say.
An operator that files no time at all in a period — neither grade — is still left out of the national figure and listed by name beside it, with the service it runs, and every national figure carries the share of the country's stop events it speaks for. That is now rare.
Counting both grades is what makes the history readable at all: it takes the measurement axis from 75.9–94.3% across the archive to 97.0–98.5%, and it moves the published punctuality figure by at most 0.45 points.
7. What we do not claim
- We do not publish a punctuality figure for service we did not observe. Where a timetable promises runs that never appear in the data, we report that as its own category — scheduled service not reported — and never fold it into cancellations or into the punctuality rate.
- We do not say an operator publishes nothing unless it does. An operator whose data we cannot connect to a line is a limitation of our matching, not a fact about them, and the two are rendered differently. Nor do we penalise an operator for the hardware at its stops: a stop with no arrival detection files a forecast, we count it, and we say what share of a figure fired an event rather than quietly dropping the ones that did not.
- We do not present a forecast as a confirmed arrival. Both are counted and the split is always available, because two things depend on it: a vehicle can pass a stop without an undetected system noticing, and where no forecast exists the value sent is the timetable — which we never score.
- We do not rank entities below the coverage threshold. Operators with the thinnest telemetry would otherwise appear the most reliable, which makes a ranking not merely imprecise but inverted.
- We do not measure across a clock change. Switzerland's civil time jumps twice a year, and the difference between two wall-clock times spanning that jump is not elapsed time. Those stop events are excluded and counted, not computed. In one recent October that was 9,203 events out of 65 million.
- We do not adjust, smooth, or exclude outliers. A delay outside −1 hour to +12 hours is classified as implausible and reported as such, rather than dropped.
- Per-line figures for rail before 2025 rest on derived connections, not published ones. They are computed the same way and checked the same way, but the connection underneath them was inferred from the journeys themselves.
8. Where the data comes from
Both datasets are published by opentransportdata.swiss under an open licence. The actual-data files and the timetable snapshots are the complete input; there is no private feed, no operator agreement, and no manual adjustment anywhere in the pipeline.
If a figure on this site looks wrong to you — particularly if you are the operator it describes — the matching step and the coverage for every number are recorded, and we would rather correct it than defend it.