Claims lab · the short version

Can code settle routine insurance claims on its own, with a narrow model answering?

Tested on fresh synthetic claims; not validated in real use

Figures last updated , from the published evidence. What changed

In the lab's baseline, a person reviews every claim. Here code decides, and Jev, a narrow model, only answers questions.

  1. A claim arrives.
  2. Jev answers fourteen questions, each with a probability. It never decides.
  3. Code sends large or suspicious claims to a person. For the rest, the bigger the claim, the more certain the answers must be.
  4. Code pays, asks for documents, or sends the claim to a person. Nothing is denied automatically.

The result

One test declared in advance: the version in shadow mode (no automatic payment over $20,000), on 3,000 fresh synthetic claims drawn after the limit. It is one test; every set pooled is below.

Paid with no person1,4861,486 of 3,000 (49.5%; range 47.7% to 51.3%)
Asked for documents, with no person127127 of 3,000 (4.2%; range 3.6% to 5.0%): a next step, not a resolved claim
Wrong decisions1in 1,613 decided with no person: 99.94% right (99.65% to 99.99%), judged on the bottom against a 99.5% bar
Net per thousand claims+$4,059on this test, drawn after the limit (range −$1,729 to +$9,519). The range crosses zero: one set cannot show the money is positive.

How the test was declared, and how to check it

Beside it: every set, pooled

All seven sets, 10,000 fresh claims, the test above included: 49.4% paid with no person, 4.3% asked for documents, 4 wrong (99.93% right, 99.81% to 99.97%). Net per thousand +$4,887 (+$1,861 to +$7,951), looking back: the payment limit came after five of the sets. On the other set drawn after it: −$1,692 (−$12,612 to +$8,120).

Where a hundred fresh claims wentEvery set pooled, the test above included: each dot is a hundredth of the 10,000 fresh claims, rounded.
  • Paid with no person: 49.4%
  • Asked for documents: 4.3%. Not resolved yet.
  • To a person: not certain enough for the amount, or over the payment limit: about 25%
  • Must reach a person by rule (the largest claims, suspected fraud or manipulation): 21.2%

So at most 78.8% could ever be settled with no person.

Every set of fresh claims

The first version's testThe corrected version's testIts larger confirmationAdding a fifteenth questionAsking the customer firstAsking first, with a payment limitThe test declared in advanceNoneAll

Hatched: the set had a wrong decision.

Fresh claims fromClaimsPaid with no personWrong
The first version's test1,000499 (49.9%)1
The corrected version's test1,000507 (50.7%)0
Its larger confirmation2,0001,015 (50.7%)1
Adding a fifteenth question1,000482 (48.2%)0
Asking the customer first1,000474 (47.4%)0
Asking first, with a payment limit1,000473 (47.3%)1
The test declared in advance3,0001,486 (49.5%)1
Each set was drawn for the test named; all are scored with the version in shadow mode.

It looked ready in development, then failed on fresh claims. A corrected version passed three times: its own test, a stricter confirmation, and a test declared in advance with the payment limit. Every test is published, failures too: three passed, four failed. Every test, in full · Each wrong decision, step by step

What this shows

  • On a test declared in advance, code with a narrow model paid about five in ten of these synthetic claims with no person, and was right on 99.94% of those it decided.

What it doesn't

  • That it works on real claims: these are synthetic, with answers the lab set.
  • That large planted frauds are caught: the model misreads a few, so nothing over the payment limit is paid automatically.
  • That subtler manipulation is caught: every attempt came from five fixed test phrasings.
  • That staff time is saved: the minutes per task are assumed, not timed.
  • That it saves money in general: this is one test on synthetic claims, and the pooled figure looks back.

Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab's owner paid for every model call, Jev's and the language models', personally. The lab's owner designed, ran and judged every test, working with AI coding assistants; nothing here has yet been replicated independently.