Ledger Kernel · research-stack
MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
seq 1Note2026-09-23

A model that cannot write, put to work measuring.

We reached for TypeSafe's Jev as a cheap extractor. It cannot produce text at all — only a probability, a choice among options you supply, or a score on a scale you define.

the same page, the same question · two kinds of answer a writing model answers in words Jev returns none of this — it has no text output no sentence, no quote, no claim, at any price a decision model answers in types probability one number in 0 … 1 choice over options you supply a b c score on a scale you define output is free because there is no output

So it could never have been our extractor, at any price. The useful question was not what it can write but what decisions we are making blind, at a scale nobody would pay a writing model to touch.

seq 2Unitthe shape

One request carries one page against every question.

one request · one page · every question you have about it one stored page already on disk question 1 question 2 question 3 question 4 … question n every question you have about that page one request page read once $0.042 / M in output: free on the page not on the page on the page not on the page on the page each answer is a probability 0 at the left tick, 1 at the right 500 questions × every stored page of ten runs = 343 requests · 3,063,489 input tokens · $0.1287 a page is paid for once, however many things you need to know about it

A judgment costs about a thousandth of a cent, and the page is read once however many things you need to know about it. That shape is why 500 questions over a whole corpus came to thirteen cents.

seq 3Doctrinecalibration

First, prove the instrument.

Before trusting a number from it we asked it about facts that do not exist: eight real rubric items and six fabricated ones of identical shape, put to the same 26 pages.

the control · the same 26 pages · probability the page states the fact threshold 0.70 0.84 – 0.98 8 real rubric facts 0.01 – 0.10 6 fabricated facts identical shape no overlap — 0.74 of the scale 0 0.25 0.50 0.75 1.00 probability that the page states the fact no overlap, so a 0.70 threshold has no false positives on this control an instrument you have not calibrated is an opinion with a decimal point

Nothing in between, so the threshold is not a judgement call. A cheap instrument that agrees with you is worse than no instrument — this same model had already been confidently wrong twelve times out of twelve on a reasoning task.

seq 4Argumentthe finding

Where the score was actually going.

Every information-recall item of ten external benchmark tasks — 500 of them, written by other people — put to every page those runs had already stored.

500 information-recall items · where the answer actually was 343 typed judgments over 3.1 M tokens of our own stored pages · $0.1287 207 293 500 rubric facts already downloaded never made into a claim never retrieved the run did not read it in benchmark points 0 100 16.42 what it scores GOLD-EXT, the clean ten 30.8 sitting unread 207 items, already on disk more than the whole score is sitting in pages already paid for, unread and it killed the fix we were about to build — the window budget we blamed accounts for almost none of it

207 facts were on disk and never became a claim: 30.8 points of a 100-point benchmark against the 16.42 the tool scores. More than the whole score, already paid for, unread.

then the sharper one

We asked whether any single passage carried a complete fact.

one rubric fact · four parts · no single passage carries them all one rubric fact as the rubric writes it the study's name the country the design the sample size four parts, one fact passage A one window name country design sample size not complete passage B one window name country design sample size not complete passage C one window name country design sample size not complete asked for the whole fact in one passage 0.48 – 0.64 0.04 real rubric facts controls present in pieces, complete nowhere — so the fix is assembly, not more reading the same 0-to-1 scale as the control · EXTRACT-1 and ASSEMBLE-1 are the next packages, not yet built

Present in pieces, complete nowhere. So the problem was never reading more — it is assembly, joining facts across passages into records, and that killed the fix we were about to build.

seq 5Budgetwhat it cost

The price of knowing.

what the measurement cost · every judgment made for this page $0.1287 the funnel 500 items × every stored page $0.0140 the control real facts against fabricated $0.0643 complete or partial 680 overlapping windows $0.0001 duplicates 18 claim pairs $0.27 total the four above are $0.2071 $0.00 $0.07 $0.14 $0.21 $0.28 before · a full benchmark pass a few dollars · a day after · the funnel $0.13 · twenty minutes a hypothesis can be killed before it is built — the return is a cheaper way to find out we were wrong

The loop matters more than the total. Diagnosing a change used to cost a full benchmark pass; the funnel is thirteen cents and twenty minutes, so a hypothesis can be killed before it is built. That is the whole return — not a better answer, a cheaper way to find out we were wrong.

next: the MVP — what the tool scores, and how that is measured — or the roadmap.