On a test declared in advance, code and a narrow decision model paid 49.5% of fresh synthetic claims with no person and asked for documents on 4.2%, right on 99.94% of the claims it decided.
The test: the version in shadow mode, with its payment limit, on 3,000 fresh claims drawn after the limit, declared before it ran. It is one test; every other test and every set pooled are below. A decision model, Jev, answers narrow questions, code makes the decision, people take the rest. The claims are synthetic, and nothing here is validated in real use.
Figures last updated , from the published evidence. What changed
49.5%paid with no person (1,486 of 3,000 fresh claims; range 47.7% to 51.3%), by the version in shadow mode, which pays nothing automatically over $20,000
4.2%asked for documents, with no person (127 of 3,000; range 3.6% to 5.0%): a request is a next step, not a resolved claim, and is never added to the paid share
1wrong decision (1 paid, 0 asked for documents) in the 1,613 decided with no person: 99.94% right (range 99.65% to 99.99%), judged on the bottom of the range against a 99.5% bar
+$4,059net per thousand claims on this test, drawn after the limit (range −$1,729 to +$9,519). Its range crosses zero: one set cannot show the money is positive.
Every set the version can be scored on, the test above included. This pooled money figure looks back.
49.4%paid with no person (4,936 of 10,000 fresh claims), pooled over seven sets, by the version in shadow mode, which pays nothing automatically over $20,000
4.3%asked for documents, with no person (432 of 10,000): a request is a next step, not a resolved claim, and is never added to the paid share (each share is rounded on its own)
4wrong decisions (3 paid, 1 asked for documents) in the 5,368 decided with no person: 99.93% right (range 99.81% to 99.97%), against a 99.5% bar
+$4,887net per thousand claims, pooled: a retrospective analysis (range +$1,861 to +$7,951)
The payment limit was set after W4's test made a wrong payment of $49,999, so the pooled money figure looks back: five of the seven sets were drawn before the limit existed. On sets drawn after it was set, the version in shadow mode nets −$1,692 per thousand claims (−$12,612 to +$8,120) on the first, and +$4,059 (−$1,729 to +$9,519) on a test declared in advance. Without the limit the pooled range crosses zero.
Every test on fresh claims
Every test on fresh claims: 3 passed and 4 failed. The figures are in each row.
The design that passed: W3, stricter as amounts rise
Margins: accuracy 99.08% against the 99.5% bar, missed by 0.42 points; it needed 3 fewer wrong decisions to pass. Net −$6,363 per thousand claims against $0, missed by $6,363; every claim analysed.
Margins: accuracy 100.00% against the 99.5% bar, cleared by 0.50 points; the bar would still have absorbed 2 more wrong decisions. Net +$6,790 per thousand claims against $0, cleared by $6,790; every claim analysed.
W3, second attempt: confirmation2,000 more, judged more strictly
Margins: the bottom of the accuracy range 99.53% against the 99.5% bar, cleared by 0.03 points; one more wrong decision would have failed it. Net +$12,613 per thousand claims against $0, cleared by $12,613; every claim analysed.
W3, second attempt: declared in advance, with the payment limit3,000 fresh claims, drawn after the limit
Margins: the bottom of the accuracy range 99.65% against the 99.5% bar, cleared by 0.15 points; the bar would still have absorbed 1 more wrong decision. Net +$4,059 per thousand claims against $0, cleared by $4,059; every claim analysed.
Attempts to go further
W3 + a fifteenth question1,000 fresh claims
65.6% settled alone (paid and asked together) · 4 wrong
Failed99.39% right: below the bar
Margins: accuracy 99.39% against the 99.5% bar, missed by 0.11 points; it needed 1 fewer wrong decision to pass. Net +$13,196 per thousand claims against $0, cleared by $13,196; every claim analysed.
W4: ask the customer1,000 fresh claims
59.9% settled alone (paid and asked together) · 1 wrong
Failednet −$43,069 per 1,000: lost money
Margins: accuracy 99.83% against the 99.5% bar, cleared by 0.33 points; the bar would still have absorbed 1 more wrong decision. Net −$43,069 per thousand claims against $0, missed by $43,069; only 999 of 1,000 claims analysed, short of the complete run the rule requires.
W4 + a payment limit1,000 fresh claims
52.2% settled alone (paid and asked together) · 1 wrong
Failednet −$3,372 per 1,000: lost money
Margins: accuracy 99.81% against the 99.5% bar, cleared by 0.31 points; the bar would still have absorbed 1 more wrong decision. Net −$3,372 per thousand claims against $0, missed by $3,372; only 905 of 1,000 claims analysed, short of the complete run the rule requires.
Share of all claims settled with no person, paid and asked for documents together, on claims each version had never seen. A test passes only if enough of those decisions are right and the version saves money after paying for its mistakes. Where a row shows paid and asked apart, they were recomputed from the stored answers and checked against the recorded result. The accuracy bar is 99.5% right. Each track runs to the ceiling, 78.4%: the most that could ever be settled alone.
1The problem
In the lab's baseline, a person reviews every claim, even the routine ones.
Today an adjuster reads every claim, checks the policy by hand and decides. That is about 122.0 min of staff time per claim and about 5.0 days of waiting, most of it in a queue, though most claims are routine.
Today. Of 6 steps, people do 5, code 1, the decision model none.
PeopleDecision modelCodeWhere claims end up
1Code: Claim arrives
2People: Waits in a queue
3People: Adjuster reads everything
4People: Adjuster checks the policy by hand
5People: Adjuster decides
6People: Payment is set up
Where the 1,000 claims end up:
Where claims end up
Step 6, Payment is set up → All 1,000 decided by a person, at step 6
The question
Can a narrow, cheap decision model settle routine claims with no person, if it only answers questions and code makes every decision?
What counts as success
At least 99.5% of the decisions it makes alone are right
Cheaper than doing it by hand, after paying for every mistake
Holds on claims it has never seen
As many claims settled alone as possible: at most 78.4% can be, because the rest must reach a person by rule
What stays with people
Every claim over the large-claim limit
Every denial: none is automatic
Signs of fraud, and any attempt to manipulate the model
2How it's built
The decision model answers questions. Code makes every decision.
The decision model is never asked what to do with a claim. It answers fourteen narrow questions, each with a probability: is the claim complete, do the photos match, how strong are the fraud signs. Code turns the answers into a decision: pay, ask the customer for what is missing, or send the claim to a person.
The decision model here is Jev, TypeSafe's first System One model, trained with RLCD: reinforcement learning for calibrated decisions. Jev is TypeSafe's model; the lab calls it through OpenRouter and changes nothing about it. A language model earns its cost only where it does what the decision model cannot: write, read a claim's documents, or check a reply against them. Changing the models, and the process around them, is how the lab works to raise accuracy and speed.
With no model, on rules alone, the same W3 method settles 4.7% of the development claims with no person and nets +$3,290 per thousand claims; with the decision model, 62.6% and +$8,809.
The design: Jev answers, code decides, people review. Of 11 steps, code does 8, the decision model 1, people 2.
PeopleDecision modelCodeWhere claims end up
1Code: Claim arrives
2Decision model: Jev answers narrow questions
3Code, a fixed rule: Over $50,000?
4Code, a fixed rule: Manipulation attempt?
5Code, a fixed rule: Fraud signs: how strong?
6Code, a fixed rule: Policy in force?
7Code: Certain enough for the amount?
8Code: Over $20,000?
9People: Quick confirm, requests only
10Code: Pay, or ask for what is missing
11People: Audit, afterwards
Where claims go:
Where claims end up
Step 3, Over $50,000? → a person, with the analysis attached, leaving at step 3
Step 4, Manipulation attempt? → quarantine, leaving at step 4
Step 5, Fraud signs: how strong? → a light check or an investigation, leaving at step 5
Step 6, Policy in force? → a person, who confirms any denial, leaving at step 6
Step 7, Certain enough for the amount? → a standard review by a person, leaving at step 7
Step 8, Over $20,000? → a person approves the payment, leaving at step 8
Step 10, Pay, or ask for what is missing → paid, or asked for documents, with no person, at step 10
fixed rule Steps marked as fixed rules run before anything else. No answer from the decision model can change them. The lab publishes its exact thresholds so the evidence can be checked. An insurer running this design would keep its operating thresholds confidential, shared with its regulator, and add random audits, because published cut-offs invite claims tailored to pass them.
Where the claims go under the recommended designEach bar is every claim counted, by the route it took. The figures are in the table below.
Development claims, each scored by settings fitted on the others1,000 claims, with the payment limit
Each fold's cut-off chosen on claims its models had not seen: the corrected way, as in the version that passed.
The payment limit sends 56 of these to a person.
Every set of fresh claims, pooled10,000 claims, with the payment limit
The payment limit sends 507 of these to a person.
Route
Development claims
Fresh claims, pooled
Paid with no person
487 (48.7%)
4,936 (49.4%)
Asked for documents, with no person
50 (5.0%)
432 (4.3%)
A quick confirmation by a person, for requests only
36 (3.6%)
296 (3.0%)
A light fraud check by a person
78 (7.8%)
735 (7.3%)
A person: not certain enough for the amount
107 (10.7%)
1,238 (12.4%)
A person, by a fixed rule or the payment limit
242 (24.2%)
2,363 (23.6%)
The recommended design on the fresh claims: the frozen W3 with the payment limit decides which claims are paid, asked for more or sent to a person, exactly as the fresh figures are scored. The quick-confirmation floor and the fraud grade were fitted once on the 1,000 evaluation claims and applied unchanged to the fresh claims, which they were never fitted on. Both were fitted with the payment limit, as the design runs.
Decision model: Jev (TypeSafe, System One) beside the language model
Measure
Decision model: Jev (TypeSafe, System One)
Language model: Claude Haiku 4.5
Version used
~typesafe/jev-latest (answered by typesafe/jev-1.13-20260917)
anthropic/claude-haiku-4.5
Trained for
Calibrated decisions (RLCD)
Answers people prefer, or that can be checked (RLHF, RLVR)
What it returns
Typed values, each with a probability; nothing to parse
Text, which code must parse and check
Speed here
A median of 276 ms a call; one call in twenty took longer than 515 ms
Not measured in this lab
Cost here
$0.12 per thousand claims, all fourteen questions
$0.81 per thousand claims for a second opinion; $1.92 for a question and its reply
Its job here
Answers the fourteen questions on every claim
Writes the question to the customer and, in this lab, the customer's reply. Proposed, not built: reading a claim's documents, and checking a reply against them, with code checking what it reads
What the lab found
Its yes-or-no probabilities sit 9.6 points from how often the answer was yes, on average. Mostly too cautious: when it said 83%, the answer was yes 94% of the time
As a second vote on the same facts it added little: asked on 110 claims, it let 4 more be settled alone. Kept as a negative result. A made-up reassuring reply once got a planted fraud paid
Ten ideas. Four designs faced fresh claims, in seven tests; one passed.
Every workflow keeps the same machinery: the decision model answers questions, fixed rules run first, and a decision rule turns the answers into pay, ask or a person. Each one changes one part of it. The main chain fixed one flaw at a time until W3 held up on fresh claims.
The main chain: each one fixes the last one's flaw
W0Trust the one big answerBuilt · misses the bar
W1Combine the small answersBuilt · loses money
W2Learn how much to trust each answerBuilt · paid a fraud
W3Stricter as amounts risePassed twice
Tried after W3
W3 +A fifteenth questionFailed on fresh claims
W4Ask the customer + Language modelFailed on fresh claims
W4 +A payment limit + Language modelFailed on fresh claims
W5A second opinion + Language modelExperiments only
In the plan, not built
Plan W3Ask twiceIn the plan, not built
W6Policy changesIn the plan, not built
Passed on fresh claims
Failed on fresh claims
Built, measured in development
Experiment, never frozen
Planned, not built
Every workflow uses the decision model, Jev. Those marked + Language model also use the language model, Claude Haiku 4.5.
Why the numbers don't follow the plan. The plan named W0 to W6. Its W3, "ask twice", was never built, and the number went to "stricter as amounts rise", an idea that came from W2 paying a planted fraud. The plan's W4 was built as a frozen workflow; its W5 was tried only as experiments.
Around W3, not workflows: The quick confirmation, the graded fraud check and the audit keep W3's decision and change the human work around it. They are on the process diagram, not in this tree.
The main chain, one line each; open any for its card
W0Trust the one big answerSettle a claim alone only when the decision model confidently answers the broad question, "what do you recommend?"3.4% alone · misses the bar
Why it exists
The starting point: the policy as first specified.
One broad answer→Fixed rules→Decision rule→Pay · Ask · Person
Reads
The decision model's one recommendation, plus the evidence checks
Learned from data
No: hand-set
Certainty needed
The same for every claim
Models
Decision model only: Jev, asked once and replayed
The planted fraudsent to a personright
In development3.4% settled alone (paid and asked together) · 70.59% right · −$100.0k per 1,000 claims
The exact rule
The engine as built: after the fixed rules, act only if the recommendation is "approve" or "ask for information" at or above the confidence threshold, and the documents, photos, estimate and fraud score all pass.
Lesson. One broad answer is rarely confident, and when it is, it is often wrong. Next: stop asking the broad question.
W1Combine the small answersIgnore the recommendation. Code reads the narrow answers (complete? photos match? fraud risk?) and acts only if even the weakest one is strong enough.36.6% alone · loses money
Why it exists
W0 leaned on one broad answer the decision model is bad at.
Eight answers→Fixed rules→Weakest answer→Pay · Ask · Person
Reads
Eight narrow answers
Learned from data
No: hand-set
Certainty needed
The same for every claim
Models
Decision model only: Jev, replayed
The planted fraudsent to a personright
In development36.6% settled alone (paid and asked together) · 100.00% right · −$76.8k per 1,000 claims
The exact rule
After the fixed rules and the specialist check, the approval confidence is the weakest of the narrow answers; act only if it clears the cut-off chosen for 99.5% right on the training claims.
Lesson. Right almost every time, but far too cautious: one doubtful answer sends a good claim to a person, so it costs more staff time than it saves. Next: weigh the answers instead.
W2Learn how much to trust each answerA small statistical model, fitted on past claims, learns a weight for each of the fourteen answers and turns them into one score.57.8% alone · paid a fraud
Why it exists
W1 treated every answer as equally important, and the decision model's probabilities sit 9.6 points from the truth on average on these claims.
Fourteen answers→Fixed rules→Learned weights→Pay · Ask · Person
Reads
All fourteen answers
Learned from data
Yes: two small logistic models
Certainty needed
The same for every claim
Models
Decision model only: Jev, replayed
The planted fraudpaid automaticallywrong
In development57.8% settled alone (paid and asked together) · 99.83% right · −$39.4k per 1,000 claims
The exact rule
Two logistic models score "paying is right" and "asking for information is right"; the likelier one is proposed, and acted on alone if its score clears the cut-off.
Lesson. Accuracy is not enough: it is right almost every time, but one costly mistake, a planted fraud of $49,101, wipes out every saving. Next: ask for more certainty when more money is at stake.
W3Stricter as amounts riseThe same learned weights as W2, but above $5,000 the doubt allowed shrinks as the amount grows: at twice that amount, the doubt must be half as large.Passed twice
Why it exists
W2 needed the same certainty to pay a small claim or a large one.
Fourteen answers→Fixed rules→Weights × amount→Pay · Ask · Person
Reads
All fourteen answers, and the amount
Learned from data
Yes: the same weights as W2, frozen
Certainty needed
Rises with the amount above $5,000
Models
Decision model only: Jev, replayed; asked afresh in each test on fresh claims
The planted fraudsent to a personright
On fresh claimsThe first version failed. A corrected version passed, then passed a larger, stricter test. The tests in full.
The exact rule
For a payment, the doubt is one minus the score, times the larger of one and the amount divided by $5,000. It pays alone only if one minus the doubt clears the cut-off: 0.813 in the version that passed, chosen on claims its models had not seen. Asking for information pays nothing, so it is never scaled.
Lesson. Holds the accuracy bar on claims it had never seen; it saves money only with its payment limit. Its known weak spot: a few large planted frauds the decision model misreads.
Follow one claim
A life benefit that is a planted fraud
life claim · $49,101
The fraud W2 paid and W3 sent to a person, drawn version by version.
Development scores flatter. Only fresh claims count.
W3 was frozen and run once on a thousand claims it had never seen, drawn from a seed chosen only when the run began. It failed: it had chosen its certainty cut-off on the same claims it was scored on. The corrected version chose it on claims its models never saw, passed, then passed a larger and stricter test.
Accuracy on fresh claims against the 99.5% bar. First attempt: failed; Second attempt: passed; Its confirmation: passed. The figures are in the text below.
First attempt1,000 fresh claimsFailed
Second attempt1,000 fresh claimsPassed
Its confirmation2,000 more fresh claims, judged on the bottom of its rangePassed
Each line is the honest range of the accuracy on claims settled with no person; the dot is the measured figure, filled when the attempt passed and open when it failed.
First version
+$5,272 per 1,000 claims in development →−$6,363 on fresh claims
Corrected version
+$5,939 in development →+$6,790 on fresh claims, then +$12,613 on its confirmation
Dollars at the staff costs each version was frozen with.
Three attempts to settle more claims all failed on fresh claims.
Each passed in development, earned its test, and failed on fresh claims. The steps they added worked as designed. What failed each test was a claim the plain W3 inside it paid with confidence, or a side effect of the new input.
W3 + a fifteenth question
Asks the decision model whether the damage could predate the incident: the one kind of claim that caused most of W3's early mistakes.
65.6% settled alone (paid and asked together) · 4 wrongFailed
Wrongly paid $5,634; net +$13,196 per 1,000 claims.
The question caught what it was aimed at. But its answer also made W3 ask other customers for information they did not need, and a wrong request counts against accuracy.
W4: ask the customer
When one piece is missing, ask the customer for it, then decide again. A reply can settle a claim only up to a cap. Set for the lab: a reply settles at most $20,000.
59.9% settled alone (paid and asked together) · 1 wrongFailed
Wrongly paid $49,999; net −$43,069 per 1,000 claims.
The step that asks the customer made no wrong payment, and caught every attempt to manipulate it, against five fixed test phrasings; subtler attacks are not yet tested. The claim that failed the test was one plain W3 paid by itself.
W4 + a payment limit
Nothing over the limit is paid without a person, whichever step proposes it. Set for the lab: a reply settles at most $20,000; no automatic payment over $20,000.
52.2% settled alone (paid and asked together) · 1 wrongFailed
Wrongly paid $2,322; net −$3,372 per 1,000 claims.
The limit stopped the large mistake, but sent correct large payments to a person, which cost more staff time than it saved. A fault in the run also left some claims unasked, so its evidence is incomplete.
its share settled alone (paid and asked together) the corrected W3's, on its own test
What none of them fixed
A few large planted frauds that the decision model misreads.
A large planted fraud near the large-claim limit was paid in testing, by the confirmed W3 too. That is why the version now running in shadow mode pays nothing automatically above a set limit. A limit on automatic payments stops such a claim but costs more than it saves, and a language model reviewing every large payment, tried in development, flagged many honest ones and still let planted frauds through. Where to draw that line is a policy decision, not something the data can choose.
6The people it takes
127 people by hand. 95 with the new process.
Staff time turned into the number a business plans around. The saving comes from settling routine claims with no person and from sizing the human work to the claim. The quick confirmation is kept only for requests for more information, where a missed wrong file costs a delay, not a payment. The graded fraud check is counted with a standard review for every claim its light check clears, and as a check that misses no fraud, so its part of the saving is still an upper bound. Both at 100,000 claims a year.
127 people by hand95 with the new process
Development figures: each fold's cut-off chosen on claims its models had not seen, as in the version that passed, and the payment limit. The version that passed paid 54.9% of its fresh claims with no person and asked for documents on 4.4%; on the larger confirmation, 54.8% and 4.8%. Paid and asked apart: recomputed from the stored answers and checked against the recorded result. The headline's version paid 49.4% with no person pooled over every set, and asked for documents on 4.3%.
A narrow decision model plus code can settle a large share of these claims with no person, above the accuracy bar, on claims it never saw.
Scaling the certainty needed to the money at stake is what made that hold.
Development scores overstate. Every design that failed looked ready first.
It doesn't show
That it works on real claims. The tests on fresh claims show the result holds on new claims from the same generator, judged against answers the lab itself set; nothing yet tests it on real ones.
That the staff savings are real. Minutes per task are benchmark assumptions, not timed.
That large planted frauds are caught. One was paid in testing, by the confirmed W3 too.
That manipulation is reliably caught. Most attempts in the fresh claims were quarantined, but not all: some reached a decision with no person (see every wrong decision), and every attempt was built from five fixed test phrasings.
That it decides real claims. W3 runs in shadow mode only: it records what it would decide and changes nothing.
That it saves money without its payment limit. Pooled over the same fresh claims, its net per thousand claims without the limit ranges from −$8,139 to +$10,728: one wrong payment of $49,999 decides it.
8Questions
Questions
Was the corrected version tuned on the test it passed?
No. The fix was worked out on the development claims only. Each test draws new claims from a seed chosen when the run starts and uses them once; no test's claims are used to fit a design. One rule did follow a test's result: the payment limit, set after W4's test made a large wrong payment. That is why the pooled money figure with the limit is a retrospective analysis. A result is final for the version tested: the first failure stays on the record.
Why synthetic claims?
Because every claim then has a known right answer, so every decision can be marked right or wrong, and frauds, manipulation attempts and missing documents can be planted on purpose. The cost is realism: a pass here is where the evidence starts, not where it ends.
Why does code decide, not the model?
A model asked to decide can be talked into a decision, and cannot say how sure it should be. Asked narrow questions, its answers can be weighed, checked against fixed rules, and priced against what a mistake would cost.
What would it take to use this on real claims?
A test on real claims, with their real outcomes; human tasks timed rather than assumed; a decision on large payments; and a period in shadow mode, where the system records what it would do while people keep deciding.
Where are the raw data and the code?
Published: a summary of every version and every test on fresh claims, and the full record of eleven example claims from one run, each file byte for byte as it was exported. Every figure on this site traces to one of them. Not published: the records of the other claims, and the code, which is in a private repository. The manifest names each file and its hash.
Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab's owner paid for every model call, Jev's and the language models', personally. The lab's owner designed, ran and judged every test, working with AI coding assistants; nothing here has yet been replicated independently.