Track record
Every scored prediction this system has made, and — more importantly — the ones it missed. Published because an unverifiable forecast is worse than none.
19 caught. 27 missed.
41.3%of confirmed events had a signal in front of them
Over 170 days on an independently verified corridor, 46 events occurred that Vaethra could check itself against. Nineteen had a statistical signal preceding them inside the scoring window. Twenty-seven did not.
That number is on the front of this page rather than buried, because the alternative — quoting precision, or the six clean hits below, and leaving the misses to be discovered — is how forecasting products earn the reputation they have. A system that only publishes what it caught is indistinguishable from one that caught nothing.
The same figures are in the product and in the API:
/api/v1/accuracy returns missed next to caught for exactly this reason.
What it caught
Six scored cases, every one with a lead time and a ground-truth link. The evidence is geolocated third-party confirmation — not our own record agreeing with itself.
| Signal fired | Family | Severity | Lead | Event | Ground truth |
|---|---|---|---|---|---|
| 2026-03-16 | capacity | 83 | 15d | 2026-03-31 | geolocated ↗ |
| 2026-03-17 | capacity | 84 | 14d | 2026-03-31 | geolocated ↗ |
| 2026-03-18 | capacity | 85 | 13d | 2026-03-31 | geolocated ↗ |
| 2026-05-21 | transits | 82 | 2d | 2026-05-23 | geolocated ↗ |
| 2026-05-22 | transits | 82 | 1d | 2026-05-23 | geolocated ↗ |
| 2026-05-26 | capacity | 81 | 1d | 2026-05-27 | geolocated ↗ |
Lead time is the gap between the signal and the confirmed event. Median 13.5 days for capacity anomalies, 1.5 for transit anomalies — the slower series warn earlier, which is what you would expect and is worth saying because it is the useful part.
What it missed, and why that is the honest half
Twenty-seven confirmed events arrived with no signal in front of them. They are not listed individually because a miss has no signal to point at — the record of a miss is the event, and the absence.
The reasons are ordinary and worth stating plainly:
- The instrument does not watch it. An event with no measurable precursor in traffic, capacity or transits cannot be anticipated by a system that reads traffic, capacity and transits.
- The precursor was inside the noise. A real deviation smaller than the series' own variance is indistinguishable from an ordinary week until afterwards.
- The precursor was there and arrived late. Some instruments publish days behind, so the signal exists but not in time to be a warning.
None of these are fixed by tuning a threshold. Lowering the bar to catch more of the 27 raises the false-positive count, and the harness scores that too.
The harness said the test was uninformative. We publish that too.
The stored run's own conclusion, verbatim:
Uninformative — the base rate is saturated: flagging every single day already scores 96.5% precision, so no signal can demonstrate skill here. This test cannot validate or refute the system. Pick a region or event class where confirmed events are rarer.
On a corridor where something happens nearly every day, a model that says "something will happen" is right nearly every day. Precision is meaningless there, and any product quoting 96% from a setup like this is quoting the base rate and calling it skill.
So the number we lead with is recall against ground truth, and the verdict that the test could not prove skill stays attached to it. The next corridors are being chosen for rarer events, which is the only way the question gets answered.
Nothing here requires trusting us
Every case, both null baselines, per-family precision, the lead-time histogram, and the leakage checks.
How a signal is scored, what counts as ground truth, and what the null baselines are.