MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
seq 1Note2026-09-23
A model that cannot write, put to work measuring.
We reached for TypeSafe's Jev as a cheap extractor. It cannot produce text at all — only a
probability, a choice among options you supply, or a score on a scale you define.
So it could never have been our extractor, at any price. The useful question was not what it
can write but what decisions we are making blind, at a scale nobody would pay a writing model to touch.
seq 2Unitthe shape
One request carries one page against every question.
A judgment costs about a thousandth of a cent, and the page is read once however many things
you need to know about it. That shape is why 500 questions over a whole corpus came to thirteen cents.
seq 3Doctrinecalibration
First, prove the instrument.
Before trusting a number from it we asked it about facts that do not exist: eight real
rubric items and six fabricated ones of identical shape, put to the same 26 pages.
Nothing in between, so the threshold is not a judgement call. A cheap instrument that agrees
with you is worse than no instrument — this same model had already been confidently wrong twelve times out
of twelve on a reasoning task.
seq 4Argumentthe finding
Where the score was actually going.
Every information-recall item of ten external benchmark tasks — 500 of them, written by
other people — put to every page those runs had already stored.
207 facts were on disk and never became a claim: 30.8 points of a 100-point benchmark
against the 16.42 the tool scores. More than the whole score, already paid for, unread.
then the sharper one
We asked whether any single passage carried a complete fact.
Present in pieces, complete nowhere. So the problem was never reading more — it is
assembly, joining facts across passages into records, and that killed the fix we were about to build.
seq 5Budgetwhat it cost
The price of knowing.
The loop matters more than the total. Diagnosing a change used to cost a full benchmark pass;
the funnel is thirteen cents and twenty minutes, so a hypothesis can be killed before it is built. That is
the whole return — not a better answer, a cheaper way to find out we were wrong.
next: the MVP — what the tool scores, and how that is measured — or the roadmap.