Ledger Kernel · research-stack
POC · measured 2026-09-16 · every number is from the project's records, listed at the end
seq 1Evalbench.result

Measured, not asserted.

Eleven micro-benchmarks, eight gate checks and an eight-question side-by-side, every value a bench.result event whose verdict is recomputed at render time; the gate document is generated, never edited. Seven of eight checks pass. Check eight fails, and stays failed here until the gate is re-run over FIX-590.

seq 2Evalbench.result · the line

The bench line: 47 gated rows, 0 FAIL.

K-1
0.34–1.70 ms
admission p99 per kind · target ≤ 2 ms
PASS
K-2
5,375–5,648 /s
sustained events · target ≥ 500
PASS
K-3
29 / 29 · 4.17 s
rebuild hash-equal @ 100k events · < 60 s
PASS
K-4
0.09 / 0.32 / 1.78 ms
anchor exact / nfc / folded @ 315 KB
PASS
K-5
37.3 MB
peak memory, deep run @ concurrency 8 · ≤ 60
PASS
K-6
0.002–0.009 ms
guard SQL p50 @ 10k–100k events · ≤ 10
PASS
A-6
11.8 ms
fold @ 1,000 claims, 50 % density · ≤ 50
PASS
U-4
78 / 78
replay served from the ledger · 0 calls
PASS
O-1
0 repeats
20 random kill -9 · report hash equal
PASS
G-2
0 over-cap
demand 400 vs cap 100 · 100 admitted, 30 refused
PASS
U-2
11 / 11
first-try valid, live, n = 20 · gated at n ≥ 200
RECORDED
targets from TESTPLAN §1.4 · release profile · mock model server · every verdict recomputed from value and target when the report renders

Measured 2026-09-16 on a MacBook Pro (macOS aarch64, 10 cpus, 32 GB), release profile, commit 79c3370, behind a load waiter so no sibling build could land mid-line. No gated row misses its POC target. U-2's live row is recorded from the gate ledger and not gated until n ≥ 200.docs/BENCH-POC.md · U-2 row: docs/GATE-POC.md (hits 11, total 11, memo_hits 9)

seq 3Evalgate · SPEC §8

Eight checks. Seven pass. The eighth is drawn.

1 PASS 8/8 · 67 calls gold-8 headless 2 PASS 0 repeats kill -9 → resume 3 PASS 73/73 · equal replay · rebuild 4 PASS 0 mismatches anchor re-read 5 PASS 4 pairs · live reject → repair 6 PASS p95 36.5 s wall p95 ≤ 180 s 7 PASS 0 FAIL of 42 bench line 8 FAIL 1 of 13 ≠ replay == rebuild #590 · a live run at concurrency 6 cited a page a sibling task had fetched replayed at concurrency 1 the same output is rejected: “A1: url not fetched in run” FIX-590 packaged: anchors resolve against the task's own retrievals · re-run at $0

Each check is a bench.result row in the gate ledger; the gate document is generated from them. Check eight fails: one live run at concurrency 6 cited a page a sibling task had fetched, and the same output is rejected when replayed at concurrency 1. FIX-590 is packaged against exactly that; the record on this page is the one measured before it.docs/GATE-POC.md (measured to 2026-09-16T16:12:59Z, commits 24bee4e and e41f64a) · docs/PACKAGES.json FIX-590

#check (SPEC §8)recordedtargetverdict
18/8 gold-8 headless at the baseline’s terminal statematched 8 of 8 · 67 model calls== 1PASS
2kill -9 ×N → resume with 0 repeated calls, equal reportresume_repeats 0 · report hash equal== 0 · == 1PASS
3replay → 0 calls, 0 reserves, identical report; rebuild equal73 / 73 served from the ledger · 0 calls · 29 tables hash-equal in 3.553 s== 1 · < 60 sPASS
40 anchors failing the byte re-read, 0 at a non-ok fetch466 anchors · 451 resolved · 451 confirmed · 0 mismatches · 0 at a non-ok fetch== 0PASS
5≥ 1 unit.rejected followed by unit.repaired in a real run4 pairs rejected → repaired · reason invariant≥ 1PASS
6standard-depth wall p95 ≤ 180 s over 3 live runswall p95 36.518 s · max $0.057 · 7 calls, 22 cachedp95 ≤ 180 sPASS
7the bench line at its POC targets75 rows · 42 gated · 0 FAILper benchPASS
8claim_status after replay == rebuild; 0 dangling refs13 replays, 1 unequal (d7f8f21a…) · 6,485 refs, 0 dangling== 1 · == 0FAIL

docs/GATE-POC.md “The checks” — the recorded numbers column, condensed; the row ids and evidence fields are in the document

seq 4Decisionfeedback · actor owner

Eight questions, one judge: 5 of 8.

browser agents at scale deep $0.72 $0.20 92 → 12 calls NOT-WORSE liveness checks standard $0.87 $0.19 117 → 11 calls WORSE RLS as a backstop scout $0.11 $0.08 16 → 4 calls NOT-WORSE RLS super- user, practice standard $0.94 $0.17 122 → 10 calls NOT-WORSE idempotency tutorial standard $0.71 $0.15 94 → 9 calls WORSE retries and backoff guide standard $0.76 $0.16 93 → 10 calls WORSE observability and alerting standard $0.72 $0.16 99 → 9 calls no report NOT-WORSE same question re-run standard $0.14 $0.04 23 → 2 calls no report NOT-WORSE old tool · its own estimate at batch rates (real ≈ 2×) kernel · settled in the ledger, live better 0 · not-worse 5 · worse 3 → 5 of 8 · the bar was 6 · recorded as eight Decision.feedback events, actor owner

The old tool's report beside the kernel's for each of the eight benchmark questions, judged with the method written down first. Not-worse on five, worse on three; the bar was six. The kernel is far cheaper partly because it drives no browser — and that is where it loses.docs/E1-2026-09-16.md · bench/sets/gold-8/questions.json (old calls, cost, state) · kernel meta lines: Gold-8 Verdicts

The eight verdicts, after the self-check

  • Q1 · not-worse55/45 toward worse. Kernel structurally stronger — adversarial verification, honest medium confidence, no answer-vs-refutation contradiction — but never states the $0.02/browser-hour rate although it read the pricing page, gives no break-even threshold, and dates Browserbase agent runs to late 2026. Old answers decisively but contradicts its own refutation and lists one source as both supporting and disconfirming. Different failure modes, roughly even.headless browser TCO · deep
  • Q2 · worseThe question asks for the specific passive and active checks. Old maps them from primary vendor docs (FaceTec, Yoti, Jumio, Persona). Kernel answers with probabilistic ML ensembles, never addresses head pose, 3D SfM, frequency-domain or rPPG, reuses a fingerprint quote for a face claim, and its sources are mostly marketing blogs — 38 % primary against 76 %.selfie / liveness verification
  • Q3 · not-worseArguably better. Kernel gives a correct, direct two-sentence answer and two correct surprises (owner bypass; FORCE RLS does not touch superusers or BYPASSRLS). Old’s Answer says research did not run, then lists four richer surprises and four primary sources — self-contradictory.RLS as a backstop · scout
  • Q4 · not-worseBoth strong, different coverage. Old more primary and rigorous (rls.c, four CVEs, the covert channel, LEAKPROOF, PgBouncer). Kernel adds two real pitfalls old missed but repeats the superuser bypass three times, is blog-heavy (20 % primary) and says pg_dump may silently omit data when its own quote says row_security=off errors.RLS superuser / owner, incidents and practice
  • Q5 · worseOld anchored in canonical sources (Brooker, Google SRE ch. 22, the IETF draft, Stripe, Envoy) with three confirmed adversarial checks. Kernel’s tutorial-design reading is interesting but 11 % primary; the check-then-act race central to its Answer is marked unsupported by its own verification; confidence high anyway.idempotency and safe retries tutorial
  • Q6 · worseOld is a complete reference: jitter, per-try vs route timeouts, gRPC throttling, retry budgets, retriability, Retry-After. Kernel covers jitter and the 409/422 state machine from primary sources but omits retry budgets, timeouts, gRPC and retriability, says the server must return 422/409 while its own quotes say SHOULD, and pads sources with four versions of the same draft.idempotency retries backoff guide
  • Q7 · not-worseOld produced nothing (OpenRouter 402). Kernel delivers a coherent thesis — SLO multi-burn-rate alerting, high-cardinality events, static thresholds as the noise source — with one confirmed adversarial check, but has two authoritative sources among ~20 vendor blogs and never reaches pipeline-specific observability. Not-worse than nothing; not a clear pass against a good answer.observability and alerting
  • Q8 · not-worseSame claims and sources as Q7 served from cache (2 calls, $0.04), reordered findings, a slightly different Answer paragraph; substance and gaps identical. Verdict matches Q7.same question, cached re-run

The first draft had Q7 and Q8 as better; the self-check moved them to not-worse because the question's own bar is authoritative sources. The verdicts are recorded as eight Decision.feedback events, actor owner, in the gate ledger; the document is the reasoning they cite.docs/E1-2026-09-16.md — Verdicts, Tally

seq 5Fetchfetch.done · status failed

Reasoning ahead, sourcing behind.

primary-source share per question · old (Q1–Q6) vs kernel (Q1–Q8) Q1 47 % 36 % Q2 76 % 38 % Q3 100 % 75 % Q4 100 % 20 % Q5 92 % 11 % Q6 85 % 42 % Q7 5 % Q8 5 % of 179 fetch.done rows in the gate ledger stackoverflow.com 7 / 7 failed · 403 challenge reddit.com 3 / 3 failed · robots_disallowed yoti.com 3 not ok browser-use.com/pricing ok · 200 — the $0.02 miss is extraction, not fetching ahead: adversarial checks · could-not-establish · honest confidence · verbatim quotes behind: primary sources · locked out of the sites practitioners write on · ten defects filed → the evaluation doctrine (E7)

The judge's finding that matters more than the tally: the kernel's discipline is ahead, its source acquisition behind. Primary-source share roughly halves wherever the old tool found primary sources, because the interim fetcher is locked out of the sites practitioners write on. That layer is what the substrate replaces.docs/E1-2026-09-16.md method steps 2 and 6, the finding, the ten defects

The ten defects filed from the side-by-side

  • defect 1Pricing page fetched ok but the headline rate not extracted or used (Q1) — unit / extraction.
  • defect 2StackOverflow 403 challenge on every URL — no challenge handling (S3).
  • defect 3Reddit robots-disallowed — an owner robots-policy decision, or accept.
  • defect 4No primary-source preference in ranking / selection — doctrine / search unit.
  • defect 5Cross-modality quote reuse: anchors byte-match but do not support the claim (Q2) — E3 judge.
  • defect 6Duplicate findings within a report (Q4, Q7) — fold / report dedup.
  • defect 7Date hallucination (late 2026, Q1) — date awareness in prompts.
  • defect 8A source marked both contested and confirming (Q5) — verification consistency lint.
  • defect 9Claim stronger than its quote: must vs SHOULD (Q6) — claim-strength lint / judge.
  • defect 10Claim contradicting its own quote (pg_dump silently omits, Q4) — adversarial pass miss.

Filed as MVP findings; the method that found them becomes the first evaluation doctrine (E7), and the same eight runs are to be re-judged after S1, through E3.docs/E1-2026-09-16.md “Defects to file” · docs/PACKAGES.json E7

next: the roadmap — FIX-590, then the substrate, identity, evaluation, driven mode and budget packages of the MVP.