Claims lab · the short version
Can code settle routine insurance claims on its own, with a narrow model answering?
Tested on fresh synthetic claims; not validated in real use
Figures last updated , from the published evidence. What changed
In the lab's baseline, a person reviews every claim. Here code decides, and Jev, a narrow model, only answers questions.
- A claim arrives.
- Jev answers fourteen questions, each with a probability. It never decides.
- Code sends large or suspicious claims to a person. For the rest, the bigger the claim, the more certain the answers must be.
- Code pays, asks for documents, or sends the claim to a person. Nothing is denied automatically.
The result
One test declared in advance: the version in shadow mode (no automatic payment over $20,000), on 3,000 fresh synthetic claims drawn after the limit. It is one test; every set pooled is below.
How the test was declared, and how to check it
Beside it: every set, pooled
All seven sets, 10,000 fresh claims, the test above included: 49.4% paid with no person, 4.3% asked for documents, 4 wrong (99.93% right, 99.81% to 99.97%). Net per thousand +$4,887 (+$1,861 to +$7,951), looking back: the payment limit came after five of the sets. On the other set drawn after it: −$1,692 (−$12,612 to +$8,120).
- Paid with no person: 49.4%
- Asked for documents: 4.3%. Not resolved yet.
- To a person: not certain enough for the amount, or over the payment limit: about 25%
- Must reach a person by rule (the largest claims, suspected fraud or manipulation): 21.2%
So at most 78.8% could ever be settled with no person.
Every set of fresh claims
Hatched: the set had a wrong decision.
| Fresh claims from | Claims | Paid with no person | Wrong |
|---|---|---|---|
| The first version's test | 1,000 | 499 (49.9%) | 1 |
| The corrected version's test | 1,000 | 507 (50.7%) | 0 |
| Its larger confirmation | 2,000 | 1,015 (50.7%) | 1 |
| Adding a fifteenth question | 1,000 | 482 (48.2%) | 0 |
| Asking the customer first | 1,000 | 474 (47.4%) | 0 |
| Asking first, with a payment limit | 1,000 | 473 (47.3%) | 1 |
| The test declared in advance | 3,000 | 1,486 (49.5%) | 1 |
It looked ready in development, then failed on fresh claims. A corrected version passed three times: its own test, a stricter confirmation, and a test declared in advance with the payment limit. Every test is published, failures too: three passed, four failed. Every test, in full · Each wrong decision, step by step
What this shows
- On a test declared in advance, code with a narrow model paid about five in ten of these synthetic claims with no person, and was right on 99.94% of those it decided.
What it doesn't
- That it works on real claims: these are synthetic, with answers the lab set.
- That large planted frauds are caught: the model misreads a few, so nothing over the payment limit is paid automatically.
- That subtler manipulation is caught: every attempt came from five fixed test phrasings.
- That staff time is saved: the minutes per task are assumed, not timed.
- That it saves money in general: this is one test on synthetic claims, and the pooled figure looks back.
The full report Staff time: see the full report What changed, and when Each wrong decision, step by step
Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab's owner paid for every model call, Jev's and the language models', personally. The lab's owner designed, ran and judged every test, working with AI coding assistants; nothing here has yet been replicated independently.