Skip to contentBePünktli

How BePünktli measures punctuality


Every number on this site comes from two public datasets that were never designed to be joined to each other. This page explains what they contain, how we connect them, how often that connection fails, and what we therefore do not claim. It is written for people who did not build the system.


1. What a measurement is

A vehicle's journey is a sequence of stops. At each stop the timetable promises an arrival time and a departure time, and the operator afterwards reports what actually happened.

A delay is actual − scheduled, in seconds, signed. A train three minutes late is +180. A bus two minutes early is −120, and it stays negative: we never round an early departure up to "on time", because departing early is its own kind of failure — it means passengers who arrived on time for the published departure missed it.

Punctuality is then the share of arrivals that were less than three minutes late. Three minutes is the Swiss industry convention, and we publish the five- and fifteen-minute figures beside it so the choice of threshold is visible rather than load-bearing.


2. The two sources

istdaten — what happened

The Swiss public-transport data portal publishes a daily file of actual stop events (Ist-Daten): one row per stop of every journey operated in the country, roughly two million rows a day. We hold every day from 1 January 2022 onward. Up to 21 September 2026 that is 3,743,695,749 rows covering 226,428,936 journeys over 1,721 operating days.

That is four days short of the calendar's 1,725, and eight days are damaged in total: four were never published at all — those are the four missing from the count above — and four more were published truncated, so they are present but incomplete. All eight are permanent; no backfill exists anywhere. We count them as damaged rather than as days on which nothing ran, and any period containing one says so on the page.

GTFS — what was promised

The same publisher issues the national timetable in the GTFS format, roughly twice a week. Each publication is a complete snapshot of the timetable as it stood that day, including changes that were only decided a few days earlier — engineering closures, replacement buses, and seasonal variations.

That cadence matters. A journey has to be compared against the timetable that was in force on the day it ran, not against today's. Using a timetable four months stale costs about five percentage points of matching, which we measured directly: a period that ran on one old snapshot matched 95.5% of its journeys where properly-dated neighbours matched 98.6%.

We therefore keep the snapshots as an archive — 214 of them, back to 2022 and forward into next year's timetable — and every operating day is matched against the snapshot that was current on that date.


3. The columns we use

istdaten's columns are German. Here is what each one means, what we do with it, and where it is incomplete.

columnmeaninghow we use itwhere it is missing or awkward
BETRIEBSTAGoperating daythe unit everything is grouped bya journey that runs past midnight belongs to the day it started
FAHRT_BEZEICHNERjourney reference, the operator's id for this runthe first thing we try to match onreused across days, so it is only an identifier together with the date. 20 operators publish a different one in the timetable than their vehicles report
BETREIBER_IDoperator, e.g. 85:11 (SBB)identifies the companythe 85: prefix is Switzerland; Austrian operators begin 81:, which is why we never strip a fixed prefix
LINIEN_IDline identifieridentifies the linemeans two different things. For bus and tram it is a self-describing code like 85:801:701. For rail it is a train number, which is reassigned between timetable years
LINIEN_TEXTthe label a passenger seesdisplay, and a weak matching signallabels are not unique: five different lines are called "231"
PRODUKT_IDmode — train, bus, tram, …grouping and product scopesometimes empty, and not only for obscure services — ordinary regional trains appear with no mode. An empty mode is treated as mainline rail and kept, never filtered out
BPUICstop codeidentifies the stoptwo formats. Seven digits is a station; nine characters is that station plus a two-character platform code, which need not be numeric. Operators migrated during 2026, so ~80% of recent rows carry the longer form against 10.9% of the archive
HALTESTELLEN_NAMEstop name as the operator writes itnever used for identityseveral operators publish none at all, and one operator went from 1.2% filled to 100% filled in the course of a single month. Stop names come from a registry instead
ANKUNFTSZEIT / ABFAHRTSZEITscheduled arrival / departurethe promiseabsent at a journey's first stop (no arrival) and last stop (no departure)
AN_PROGNOSE / AB_PROGNOSEreported arrival / departurethe outcomepresent only when the status below says so
AN_PROGNOSE_STATUS / AB_PROGNOSE_STATUShow the reported time was obtaineddecides whether the row is a measurement at allsee below
FAELLT_AUS_TFcancelleda failure, counted in the denominator—
DURCHFAHRT_TFpassed through without stoppingexcluded — it is not a stop—
ZUSATZFAHRT_TFan extra, unscheduled runexcluded from rates — there is no promise to measure it againstit is matched to the timetable where possible, because a fifth of them turn out to be in it

The status column is the one that matters most

AN_PROGNOSE_STATUS takes one of four values, and they are not four grades of the same thing:

This column is the single easiest thing to get wrong, and getting it wrong inverts a ranking. A natural-looking test — "is the status not REAL?" — is true for UNBEKANNT and for empty, so writing it that way quietly scores every stop nobody reported on as punctual, and the operators with the worst telemetry come out the most reliable. We never score a stop with no time on it; it is excluded from the rate and counted in the coverage figure instead.

Those statuses decide what each stop event becomes. For a recent representative month, the national outcome of every stop event we were entitled to measure was:

eventssharescored?
an arrival event fired (REAL)62,329,33393.53%yes
a forecast was held (PROGNOSE)1,459,1332.19%yes
the forecast was the timetable556,0150.83%no — §6
cancelled528,4830.79%yes, always as late
no time at all1,761,1082.64%no
implausible value6,3200.01%no

The third row is the rule §6 sets out at length: a forecast filed at the scheduled second, in a journey carrying at least one other, is the plan copied rather than a measurement, so it is in no rate's numerator and in no rate's denominator.

So a headline punctuality figure for that month is computed over 96.51% of the stop events it could have used — and 97.71% of what it counted fired an arrival event. We print both.


4. The hard part: which timetable entry is this journey?

The two datasets describe the same vehicle movements and, for most of our history, share no reliable identifier.

The timetable's original_trip_id field — which names the journey reference each trip corresponds to — only appears in feeds published from 14 December 2024 onward. For everything before that, roughly three years of the archive, the timetable simply does not say which journey its trips are. And even where the field exists it does not solve the problem: in a recent month it matched 1,971,546 journeys while 2,052,286 had to be matched another way, because those operators publish a different reference in the timetable than their vehicles report.

So we match on structure: the shape of the journey itself.

A journey's stop pattern — which stops, in which order, at which times — is close to a fingerprint. Two different trains rarely call at the same stops at the same times on the same day. We use that, in five steps, taking the first that gives exactly one answer:

stepwhat it compareshow reliable
1. Published referencethe journey reference and the date, refusing if more than one trip matches98.5% correct on operators with stable references
2. Exact patternthe full multiset of stops and times, within the same day and operator0 errors in 416,020 decisions
3. Contained patternone journey's stops contained in the other's, at least 3 stops and at least 70% overlap0 errors in 7,647 decisions
4. Ends and first departurefirst stop, last stop, first departure time9 errors in 23,210 decisions — 99.96%
5. Journey numberthe operator's own run number, then checked against the stop patternused where a legacy identifier is all there is

Two rules make this trustworthy rather than merely clever.

If more than one timetable entry fits, we record no match. Choosing the most likely candidate would be the single largest source of silent error in a system like this, and a wrong match is worse than no match: it scores one line's delay against another line's schedule, and every figure downstream inherits it. Journeys that fit two entries are marked ambiguous and excluded.

Every match records which step produced it. A number is only as good as the weakest step behind it, so the step is stored with the journey and can be audited afterwards.

How often this works

96.50% of all 225,148,610 journeys are matched to their timetable entry. By year:

20222023202420252026
94.9%95.7%96.1%97.4%98.8%

The older years are lower because they have no published reference at all and rely entirely on structure. The unmatched remainder is not a random sample — it is weighted towards journeys whose timetable had changed, which means towards disrupted service. That is why we publish the matching rate next to every figure rather than only the figure.


5. The other hard part: which line is this?

Knowing which journey a row belongs to is not the same as knowing which line. The two datasets name lines in two incompatible ways, and which one applies depends on the mode.

The timetable, meanwhile, names everything by route id such as 96-701-j23-1, where the j23 part is the timetable year.

This has a consequence worth stating plainly, because it looks like a data gap and is not one: an inventory that compares the two sides directly finds thousands of timetable routes with "no observations at all", and most of that is the mismatch rather than missing service. PostAuto's line 101 appears in our data twice — as an observed line with 1,538 runs a week, and as a route with no observations and 1,629 scheduled trips a week. It is one healthy line running at 94% of its schedule.

We resolve this by deriving the connection rather than assuming it: the journeys matched in section 4 tell us which timetable route a line's runs actually belong to. That derived connection is only accepted when five conditions hold, and it is refused otherwise:

  1. the line's journeys reach exactly one route in that timetable year;
  2. that route is reached by exactly one line, so no route's scheduled service is counted twice;
  3. there are enough journeys behind the pairing to be more than a coincidence;
  4. at most a tenth of the line's year ran somewhere the pairing does not account for, and that shortfall is recorded;
  5. the stops the line was observed serving and the stops the route is scheduled to serve actually agree.

The fifth condition is what separates associated from identical. Three quarters of accepted pairings match perfectly — every observed stop is a scheduled stop and vice versa. The ones we reject include replacement-service routes that shared a single stop with a line, and routes that turned out to be a fragment of a longer line, which would have made that line's schedule look smaller than it is.

0.68% of stop events cannot be attributed to a named line. Those rows keep everything else — station, times, delay — and count in national and per-operator figures; they simply do not appear on a line's page.


6. Coverage, and why there are three numbers

A punctuality rate is only meaningful next to the share of service it was computed over. We publish three, and the figure we show is always the worst of them, never the average:

axisquestion
daysdid the publisher release the day at all?
journeyswhat share of the day's journeys reached a timetable entry?
measurementwhat share of the stop events we were entitled to measure carry a real reading?

They are independent, and any of the three can be the binding one. A recent month had all its days, matched 98.6% of its journeys, and still measured only 94.3% of its stop events — so it publishes at 94.3%. Averaging the three would have shown 97.6% and hidden the axis that was actually thin.

Where coverage is below the threshold we consider publishable, we show the state rather than a greyed-out number. A figure computed over a tenth of a line's service is not a cautious estimate; it is a different claim, and the honest thing is not to make it.

Two grades of time, both counted, and the difference always visible

Swiss operators file a row for every stop event, and each row carries a status saying what kind of time it holds. The Swiss realisation spec for that field defines two that matter:

We count both, and we keep them apart. The reason is that the missing event is a property of the equipment at a stop, not of the journey: a stop with no door signal and no loop stays PROGNOSE however plainly the vehicle served it. Until September 2025, Geneva's Transports Publics Genevois filed 5 million stop events a month that way and not one with an event — while running one of the country's largest networks. Scoring only the confirmed grade would have measured which operators own which hardware.

We checked before deciding. Those forecasts are not the timetable restated: 99.5% of them differ from the planned time, and when TPG and Lausanne's tl each began detecting arrival events — on different dates, two years apart — the punctuality distribution did not step. Their systems were tracking vehicles well before they could confirm them.

What we do not do is hide the difference. Every figure can say what share of it fired an event, because two things are only visible to someone who can see that split: a vehicle may have driven past a stop without the system noticing, and where no forecast exists at all the value sent is the timetable — which we never score, because a timetable is not a measurement of anything.

That last one is a rule with a shape, and it is worth stating exactly, because it is the difference between a punctuality figure and a copy of the plan. The standard says that when no forecast is available the value to send is the scheduled time itself. A stop reported that way is not evidence that anything arrived on time; it is the timetable, and scored against the timetable it is perfect by construction. So a forecast equal to its scheduled time to the second is not counted as a measurement — it leaves the rate and joins the coverage figure, exactly as a stop with no time at all does.

With one exception, because a vehicle really can arrive at the scheduled second. We only apply the rule where the same journey has at least two such stops. That is not a judgement call: arrivals that did fire an event and landed exactly on the scheduled second — which are genuine by definition — are the only such stop in their journey 43–67% of the time, while forecast stops filed at the scheduled second are isolated only 8–14% of the time and typically come in runs covering a whole journey. A single coincidence is kept. A journey reported entirely as its own timetable is not.

Every figure says how much of it this removed, and the effect is small nationally and large for the few feeds it is about: a tenth of a point of punctuality and under a point of coverage across the country, against whole operator-months that stop being publishable. One city bus operator ran seven consecutive months in which essentially every arrival it filed was the scheduled second exactly, and published 100.00% punctuality at 100% coverage for each of them. It now publishes nothing for those months, which is the honest answer: we cannot say.

An operator that files no time at all in a period — neither grade — is still left out of the national figure and listed by name beside it, with the service it runs, and every national figure carries the share of the country's stop events it speaks for. That is now rare.

Counting both grades is what makes the history readable at all: it takes the measurement axis from 75.9–94.3% across the archive to 97.0–98.5%, and it moves the published punctuality figure by at most 0.45 points.


7. What we do not claim


8. Where the data comes from

Both datasets are published by opentransportdata.swiss under an open licence. The actual-data files and the timetable snapshots are the complete input; there is no private feed, no operator agreement, and no manual adjustment anywhere in the pipeline.

If a figure on this site looks wrong to you — particularly if you are the operator it describes — the matching step and the coverage for every number are recorded, and we would rather correct it than defend it.