Ledger Kernel · research-stack
POC · measured 2026-09-16 · every number is from the project's records, listed at the end
seq 1Questionpremise · load_bearing

The output is prose you have to trust.

In 2026 everyone ships deep research: model labs, search engines, agentic browsers, browser-use clones. The run is not kept, so nothing about the report can be checked.

seq 2Questionthe reader’s five questions

Five questions a reader asks. Five dead ends.

a deep-research run searches · pages · browser steps · model calls — none of it kept the report fluent prose · a list of URLs Which page said this, in those words? not recorded Run it again, exactly. a different run every time What did it cost? an estimate at batch rates: half of real This claim is wrong. nowhere to put it Did anything fail? answered anyway, confidently our own tool, Aug 2026: 16–122 calls · 157–739 s · $0.11–$0.94 estimated per question · 2 of 8 runs with no report

A run produces a report; the searches, pages, browser steps and model calls behind it are not kept. Each question a careful reader asks of the report stops at the same wall: there is no record to answer it from.web-research/DEEP_RESEARCH.md · docs/RESEARCH-POC.md §12

The five, as failure modes

  • not recorded“Which page said this, in those words?” The report carries fluent prose and a list of URLs; the page that was actually read, and the bytes the sentence rests on, are gone with the run.
  • a different run every time“Run it again, exactly.” Nothing was recorded to re-execute from, so a second run is a second opinion, not a check of the first.
  • half of real“What did it cost?” The number printed is an estimate at batch rates, about half the real spend.
  • nowhere to put it“This claim is wrong.” There is no place in the run for the reader's disagreement to be recorded against the claim it judges.
  • answered anyway“Did anything fail?” Failures are answered over, confidently.
seq 3Runour own tool · Aug 2026

Our own earlier tool was no different.

The web-research tool this project replaces — a browser-use clone driving Gemini through OpenRouter — ran the same eight questions the kernel is now measured on. It answered confidently when every browser agent had failed, and printed costs at half the real spend. Its eight runs are the baseline in the benchmark set.

#jobquestiondepthmodel callssecondsest. $terminal state
Q1…432ff2browser agents at scaledeep92418.4$0.72done
Q2…8040bbliveness checksstandard117738.8$0.87done
Q3…8d2144RLS as a backstopscout16166.2$0.11done
Q4…2a60aeRLS superuser, practicestandard122598.9$0.94done
Q5…6e2645idempotency tutorialstandard94512.2$0.71done
Q6…83c859retries and backoff guidestandard93545.5$0.76done
Q7…e683edobservability and alertingstandard99506.7$0.72error · no report
Q8…fd1e34same question, re-runstandard23157.4$0.14error · no report

16–122 calls · 157–739 s · $0.11–$0.94 estimated per question · 2 of 8 runs with no report. The dollar column is the tool's own estimate at batch rates; the scout run it priced at $0.107 cost about $0.21 in fact. The two error rows ended without a report and are still on the benchmark: the kernel's Q7 and Q8 are judged against nothing.bench/sets/gold-8/questions.json (baseline calls, seconds, estimated_usd, status) · bench/sets/gold-8/caps.toml (the batch-rate note) · docs/E1-2026-09-16.md Q7 (OpenRouter 402)

next: how it works — the whole system as composed, append-only, content-addressed ledgers.