What changed, and when

Every change to the report's headline figures, newest first: when the evidence was exported, what changed, and why. A figure changes only when the evidence is exported again, and each export is published as it was, byte for byte.

Evidence release

29 September 2026, 23:28 UTC

The headline now leads with the test declared in advance: the version with its payment limit, on fresh claims drawn after the limit, which passed. Every set pooled, now seven with that test, stays beside it as context, and its money figure still looks back. The recommended design is now priced with its payment limit and the corrected cut-off. It settles fewer claims itself and takes longer, by more than its folds vary: claims over the payment limit now go to a person, and the corrected cut-off is stricter. The earlier figures reproduced exactly before the change.

FigureBeforeAfter
Fresh claims settled with no person, paid and asked for documents together (the report's headline until it was split)53.6%53.7%
The test declared in advance: paid with no person (the report's headline since it leads)–49.5%
The test declared in advance: asked for documents with no person–4.2%
The test declared in advance: wrong decisions–1
The test declared in advance: net per thousand claims–+$4,059
Pooled fresh claims settled with no person, in shadow mode53.6%53.7%
Pooled fresh claims paid with no person (the rest of those settled were asked for documents)49.3%49.4%
Pooled fresh claims: wrong decisions34
Pooled fresh claims: net per thousand claims+$5,242+$4,887
Tests on fresh claims passed23
Tests on fresh claims run67
People needed with the recommended design9095
Recommended design: claims settled with no person62.6%53.7%
Recommended design: days to a decision1.8 days2.2 days
Recommended design: staff time per claim86.0 min91.4 min
Recommended design: net per thousand claims+$37,889+$31,561

Evidence 96f9d5e30b51

Evidence release

27 September 2026, 14:12 UTC

Each held-out test and confirmation now shows how far its recorded accuracy and net result cleared or missed its bars, and whether every claim was analysed. These margins come from the stored results under the original verdict rules. No earlier figure or verdict changed.

No headline figure changed.

Evidence 1e85a1ebdf92

Evidence release

27 September 2026, 10:51 UTC

Each test on fresh claims now shows paid and asked for documents apart, wherever the split recomputed from the stored answers matched the recorded result, and the wrong decisions are split the same way. Each of the headline's wrong decisions is described, and accuracy is drawn against the share settled. No earlier figure changed.

No headline figure changed.

Evidence 8ea8613f88b2

Site change

27 September 2026, 10:41 UTC

The short version is now the first page, at /claims/, and the full report has moved to /claims/report/. The old address of the short version leads to the new one. No figure changed.

Correction

27 September 2026, 01:32 UTC

The headline's settled share is now shown as its two parts, paid with no person and asked for documents with no person, because a request for documents does not resolve a claim. No figure changed.

With evidence 3101a494

Evidence release

27 September 2026, 01:24 UTC

The evidence now splits the claims settled with no person into those paid and those asked for documents, says what the synthetic claims are made of, and names the exact model version that answered each test. The share paid with no person is published for the first time; no earlier figure changed.

FigureBeforeAfter
Pooled fresh claims paid with no person (the rest of those settled were asked for documents)–49.3%

Evidence 3101a4944008

Correction

26 September 2026, 22:09 UTC

The pooled money figure is labelled a retrospective analysis, because the payment limit was set after W4's test; the one set drawn after the limit is shown beside it. No figure changed.

Site release ed9bcab

Evidence release

26 September 2026, 14:09 UTC

The headline now pools every set of fresh claims the passed version can be scored on, under the payment limit it runs with in shadow mode, in place of its single confirmation.

FigureBeforeAfter
Fresh claims settled with no person, paid and asked for documents together (the report's headline until it was split)59.5%53.6%
Pooled fresh claims settled with no person, in shadow mode–53.6%
Pooled fresh claims: wrong decisions–3
Pooled fresh claims: net per thousand claims–+$5,242

Evidence 43c844f45624

Evidence release

25 September 2026, 22:34 UTC

The recommended design gives every claim the light fraud check clears its standard review, and the export computes how much fraud the check must catch to be worth having.

FigureBeforeAfter
People needed with the recommended design8690
Recommended design: staff time per claim82.7 min86.0 min
Recommended design: net per thousand claims+$41,739+$37,889
Light fraud check: the share of fraud it must catch to be worth having–72.7%

Evidence e91090263ce4

Evidence release

25 September 2026, 21:38 UTC

The recommended design changed: the quick confirmation is kept for requests for more information only, so every approval a person sees gets a full review. That takes more people.

FigureBeforeAfter
People needed with the recommended design7886
Recommended design: days to a decision1.4 days1.8 days
Recommended design: staff time per claim74.8 min82.7 min
Recommended design: net per thousand claims+$50,909+$41,739

Evidence a8c8221595eb

Evidence release

25 September 2026, 19:58 UTC

The export began computing how long a standard review can take before W3 stops saving money on its confirmation.

FigureBeforeAfter
Standard review: the time at which W3 stops saving money–43.3 min

Evidence 1ed8dbc87a98

Correction

25 September 2026, 18:25 UTC

A sentence about W2 was corrected to what the calibration test shows, and the fairness check now states its limits.

No headline figure changed.

Evidence 1d8d5e9ad697

Evidence release

25 September 2026, 15:58 UTC

The evidence gained the model layer: the decision model measured on the run, its calibration included.

FigureBeforeAfter
Decision model: how far its probabilities sit from how often the answer was yes–9.6 points

Evidence 175629247405

Evidence release

25 September 2026, 13:17 UTC

The tests of the three designs built after W3 were published. All three failed.

FigureBeforeAfter
Tests on fresh claims run36

Evidence 08b3c810e1fb

Evidence release

24 September 2026, 12:48 UTC

The passed version's larger confirmation on fresh claims was published, judged on the bottom of its accuracy range.

FigureBeforeAfter
Fresh claims settled with no person, paid and asked for documents together (the report's headline until it was split)59.3%59.5%
The confirmation: claims settled with no person–59.5%
The confirmation: wrong decisions–1
Tests on fresh claims passed12
Tests on fresh claims run23

Evidence 604a57ec0329

First published

24 September 2026, 02:56 UTC

The site first published its figures: both tests of W3 on fresh claims, the first failed and the corrected one passed.

FigurePublished
Fresh claims settled with no person, paid and asked for documents together (the report's headline until it was split)59.3%
Tests on fresh claims passed1
Tests on fresh claims run2
People needed by hand, at the report's yearly volume127
People needed with the recommended design78
Recommended design: claims settled with no person62.6%
Recommended design: days to a decision1.4 days
Recommended design: staff time per claim74.8 min
Recommended design: net per thousand claims+$50,909
By hand: staff time per claim122.0 min
The new process in development: settled with no person62.6%

Evidence 18205a837f90