seq 1 Argument claim.asserted · anchors[0]
Research you can hold to account.
Every deep-research tool can search, read and write. None can be held to what it wrote. Here every claim, model call, dollar and judgment is an event you can open, replay and check. The report is prose; the reason is a ledger.
report.md · Findings
Stripe caches status codes and response
bodies for keyed requests once endpoint
execution begins, replaying even 500
server errors on subsequent retries.
high confidence · docs.stripe.com
“Stripe’s idempotency works by saving the
resulting status code and body of the first
request made for any given idempotency key,
regardless of whether it succeeds or fails.
Subsequent requests with the same key return
the same result, including 500 errors.”
A1
fetch.done · status ok · docs.stripe.com
bodies/<sha256_text>.zst · extracted text, bytes
Stripe’s idempotency works by saving the
resulting status code and body of the first
request made for any given idempotency key,
regardless of whether it succeeds or fails.
Subsequent requests with the same key return
the same result, including 500 errors.
start
end
the model cites url + quote · the kernel resolves fetch_ref and [start, end) · rungs exact → nfc → folded
GATE-4 · every resolved anchor decoded from the store and compared byte-for-byte
490 / 490 · 0 mismatches
A finding in a real report, and the page it came from. The model cited a URL and a quote; the kernel found the quote in the bytes it had fetched and stored the offsets. The POC gate re-read every resolved anchor from the store: 490 of 490 matched, 0 mismatches . The MVP kept the rule and ran it over twenty campaigns: 53 questions, 708 of 723 live claims linked into the Knowledge ledger, $8.27 in all.docs/GATE-POC.md check 4 · docs/SPEC-kernel.md §2 anchor resolution · the report: Gold-8 Verdicts, Q6
seq 2 Plan four doors
Four doors.
Question premise · load_bearing
The problem
The output of deep research is prose you have to trust. Five questions a reader asks of a report, the dead end each one hits, and our own earlier tool's numbers.
/problem →
Eval build log · 21 Aug – 18 Sep
The POC
The first half as dated entries: the eight-question baseline, three judged designs, 28 packages in five days, 47 gated bench rows, and a gate of eight checks that failed once and closed at 8 of 8.
/poc →
Run build log · 19–23 Sep
The MVP
A dated build log, newest first: what was built or measured, the number it produced, and the record it came from — including the five numbers and diagnoses we had to take back.
/mvp →
Note the instrument · 23 Sep
The measurement
A model that cannot write at all — only score, choose or rate — found 30.8 benchmark points sitting unread on pages the runs had already paid for. Every judgment behind it cost $0.27.
/jev →
Not released. The MVP runs and the console is built; the repository is private. Ask for access and we send one message when a question of yours can be run and held to account.
Request access →
between them: how it works — the ledgers as rails, append-only, kill -9, replay at $0, budget — and the roadmap , 55 of the 56 MVP packages in, extraction and assembly next.
seq 3 Eval bench.result
The headline numbers, each with its file.
EXTERNAL · 23 Sep
questions, rubrics and items written by other people; no run counted read the source its own rubric came from
GOLD-EXT · the clean ten
16.42
672 rubric items · recall 12.89 · analysis 27.16 · presentation 27.04
RECORDED
JEV · the funnel
207 / 500
recall items whose material was already on disk · 30.8 points, unread
RECORDED
JEV · spend
$0.27
every judgment of that measurement, four runs · output free
RECORDED
MVP · 19–21 Sep
rendered from the gate ledger by a program; a row it cannot give says not measured
GATE-MVP · runs
53 / 53 · $8.27
20 campaigns · 52 done · 1 gap · 531 paid calls · p50 93.6 s
MEASURED
E3 · six laws
106 / 106 · 120 / 212
L1, L3 by the ledger's rule · L2 L4 L5 L6 labelled: 6 of 53 pass all four
LABELLED
E1 · re-judge
4 of 8
better 3 · not-worse 1 · worse 4 · the bar is 6
MISSED BY TWO
POC · 16 Sep
the gate as its last pass of that day recorded it, after FIX-590
GATE · check 4
490 / 490
resolved anchors re-read byte-for-byte · 0 mismatches
PASS
GATE · 8 checks
8 of 8
check 8 failed at 16:12, FIX-590 merged, re-measured at 22:49 · in CI since
PASS
E1 · side-by-side
5 of 8
not-worse 5 · worse 3 · the bar was 6
MISSED BY ONE
The POC values are bench.result events whose verdict is recomputed from value and target when the report renders. The MVP gate is rsk gate mvp report over the X2 ledger; a re-render of the same ledger is byte-identical. The external ten are one bench.result row per rubric item. Neither document is edited by hand.docs/GOLD-EXT-CLEAN.md · docs/JEV-MEASUREMENT.md · docs/GATE-MVP.md · docs/X2-REVIEW-2026-09-20.md · docs/BENCH-POC.md · docs/GATE-POC.md · docs/E1-2026-09-16.md
Sources · every number on this site, by file
docs/GOLD-EXT-CLEAN.md — the clean ten: ten DeepResearch Bench II tasks, 672 expert-written rubric items, overall 16.42 (information recall 12.89, analysis 27.16, presentation 27.04), no run having read the source its own rubric came from; the four re-run tasks 16.76 → 17.61; the three fetches the blocklist refused live; the 52-item ceiling; the published leaderboard row and why it is a different judge.
docs/GATE-MVP.md — the MVP gate rendered from the X2 ledger: 20 campaigns, 53 questions, 53 finished (52 done, 1 gap); p50 93.6 s; 531 paid calls, 5 memo hits, $8.271496; 163 of 217 ladder-eligible failed fetches recovered, 40 refused, 14 silent; 1,984 of 1,984 rows with a policy version; L1 and L3 53 / 53 by the ledger's rule; the labels, the two judges and the E1 re-judge.
docs/X2-REVIEW-2026-09-20.md — the 53 reports read by three readers: 264 defects (65 high) in 13 root-cause clusters, each with its fix and the package it became. Its two fetch-ladder sentences quoted the gate's 20 September render, “163 of 203 … 0 silent failures”; the 21 September re-render says 217 ladder-eligible, 163 recovered, 40 refused and 14 silent, the gate is the record, and the review carries that correction since 23 September.
docs/DEPTH-1-NOTES.md — plan v2 and the cover state: 91.5 → 191.7 retrieval targets and 14.6 → 24.8 claims per task at $0.26 → $0.40, 2.4 rounds; information recall +99 %, analysis −47 %; plan v1 is still the default, promotion is E4's by score.
docs/JUDGE-1-NOTES.md and docs/GOLD-EXT-V2.md — one bench.result row per rubric item; 366 refused recall items of which 212 for the report's shape; the re-score 10.67 → 16.19 → 16.58 that showed 10.67 was our own judge misreading its own rubric; the 672 verdict rows and the partial-row gate.
docs/CONSOLE.md — the console's five screens and their render times; four adversarial repair rounds, the last of them a leased key never verified against a live provider. docs/SPEC-kernel.md — the FIX-1281 erratum [C-301], and [C-361], [C-362] for plan v2.
The Jev measurement (jev.html ) — 500 information-recall items put to the pages ten runs had already stored: 207 had their material on disk = 30.8 benchmark points against the 16.42 the tool scores; the control, 8 real facts at 0.84–0.98 against 6 fabricated ones at 0.01–0.10; $0.27 in all, metered from each response's own usage.input_tokens and the one set of figures here that is not yet a row in this repository. Model jev-1.13.0, TypeSafe System One. The record is docs/JEV-MEASUREMENT.md ; the funnel and the split are issues #1404 and #1405 .
docs/BENCH-POC.md — 47 gated rows, 0 FAIL; K-1…G-2 values; 2026-09-16, commit 79c3370, macOS aarch64.
docs/GATE-POC.md — the eight checks at commit 89e22a6 (2026-09-16T22:49:15Z): 8 PASS , Failures “None”; 498 / 490 / 0 anchors; 5 repair pairs of 7 rejected; live p95 19.364 s, max $0.031; 72 calls; replay 79 / 79 at $0; 14 replays equal to rebuild over 172 claim rows; 6,130 refs, 0 dangling; the B3 refutation clause recorded, not gated. The earlier pass at be6d0d4b (7 PASS / 1 FAIL, 1 of 13 replays unequal) is what FIX-590 fixed, merged 2026-09-16 (#673); checks 1, 4, 5 and 8 run in CI over the gold-8 tree on every merge since.
docs/PACKAGES.json — 28 POC, 62 MVP and 18 production packages; the BUILD lane of six is counted outside the MVP. EXTRACT-1 (#1404) and ASSEMBLE-1 (#1405) are filed as issues and are not in the file yet.
docs/E1-2026-09-16.md — 5 of 8; the eight verdicts; primary-source shares; 179 fetch rows, 7 / 7 and 3 / 3 failures; ten defects.
bench/sets/gold-8/questions.json · caps.toml — the old tool's calls, seconds, estimated cost and terminal state; the run caps. Gold-8 Verdicts page (cited by E1) — each kernel report's meta line.
docs/SPEC-kernel.md §1–§3, §7–§8 — event shape and id, the ledgers and invariants, anchor resolution, the fold, the plan, the interfaces (CLI, MCP tools, report), the gate. docs/DESIGN-kernel.md §0, §7.2, §13 — one writer, the invariant registry and the closed list of reason codes admission::reason is held to, resume and replay. The registry is 33 codes at the POC gate (SPEC-kernel.md §1 at commit 89e22a6: A1–A6, B1–B4, D1, Dec1, E1–E5, F1–F5, P1, Pol1, Q1–Q2, R1–R5, U1–U2) and 42 on main today — A7, D2, En1, F6–F7, P13–P14, Pol2 and Q3 have been added since; the count moves, so it is dated wherever this site prints it.
docs/RESEARCH-POC.md §8, §11–§13 — the fold's complexity, the anchor ladder, the old tool's batch-rate costs and evidence-census bug. doctrine/v1/DOCTRINE.md — the six laws. web-research/DEEP_RESEARCH.md — the old tool's method and traps. docs/ADR-crates.md , README.md — the build.
research-stack · MVP, September 2026. This site states only what the project's records say; where the record is a miss, the miss is drawn. Cut from the one-page landing, site/landing/index.html.