research-stack · research.devclusterai.com
MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
seq 1Planplan.version · derived 2026-09-23

Built, measured, next.

110 work packages in three stages — 108 of them in docs/PACKAGES.json and two filed on 23 September that are not in the file yet. The 28 POC packages are built and their gate closed at 8 of 8. Of the 62 MVP packages, 6 are the CI and build-speed lane and are counted apart; 55 of the other 56 are in, and the one still open is ONBOARD-1. 18 production packages wait behind them. The two filed last — EXTRACT-1 and ASSEMBLE-1 — went on the list because a measurement said they matter more than anything already on it. Nothing here is done by assertion: every in below is a merge commit on main, or a record named on this page that stands in for one.

seq 2Planthe derivation

What counts as done, and who says so.

docs/PACKAGES.json is the list: every package’s id, lane, title, what it is, what it needs and a done-when that is measured rather than asserted. The file has no status field, so the state on this page is derived, by one rule — a package is in when git log --merges main carries a merge of pkg/<id>. That rule finds 76 of the 108 packages in the file. Eight more are in without a merge of their own, and rather than quietly adding eight to a number, each is named here with the record that stands in for its merge.

  • K0built before the one-PR-per-package rule began on 14 Septemberstands in for the merge: direct commits, 11–12 September
  • K1built before the one-PR-per-package rule began on 14 Septemberstands in for the merge: direct commits, 12–14 September
  • E1a measurement, not a code change: eight questions side by side, 5 of 8stands in for the merge: docs/E1-2026-09-16.md
  • X1a measurement: the eight POC gate checks, 8 of 8 at commit 89e22a6stands in for the merge: docs/GATE-POC.md
  • S2an umbrella; its four parts each have a merge of their ownstands in for the merge: S2a · S2b · S2c · S2d
  • G3an umbrella; the five substrate decisions are answered in the file itselfstands in for the merge: docs/PACKAGES.json, 17 September
  • X2a measurement: 53 of 53 questions over 20 campaigns for $8.27stands in for the merge: docs/GATE-MVP.md
  • FIX-1281landed in JUDGE-1's pull request rather than one of its ownstands in for the merge: inside PR #1400

Everything else is either open — filed as an issue and being built — or not started. The counts in the figure below, in the lede above it and in every heading on this page come out of that derivation when the page is built; none of them is typed in by hand, which is the only way a roadmap and its own caption stay in agreement.

every package in docs/PACKAGES.json · one cell is one package in — a merge of pkg/<id> on main, or the named record that stands for one open — filed as an issue and being built · not started — nothing on the record yet POC 28 built and gated · poc.html MVP 56 56 packages · open: ONBOARD-1 CI and build speed 6 counted apart from the MVP · not started: BUILD, BUILD-2, BUILD-3, BUILD-4, BUILD-5 filed 23 September 2 EXTRACT-1, ASSEMBLE-1 · the measurement that justified them came first production 18 behind the MVP, none started 84 in · 3 open · 23 not started · derived at build time from docs/PACKAGES.json and git log --merges main 8 packages are in without a merge of their own; each is named in the list above this figure, not assumed

The bars are to one scale, so the MVP’s 56 and production’s 18 compare honestly. An in package is green because it passed: one merge commit, three independent readers before it landed, and a done-when that had to be demonstrated in the pull request. Open is blue and unfilled — filed, not finished.docs/PACKAGES.json · git log --merges main · the eight exceptions listed above

seq 3Runstage poc · 28 packages

POC: built, gated, one fix packaged.

All 28 POC packages are in. Their gate is the measurement, and it is on the POC page: on 16 September it ran three times, failed check 8 at 16:12, and closed at 8 of 8 at 22:49 once FIX-590 had merged (issue #640, PR #673). Checks 1, 4, 5 and 8 have run in CI over the gold-8 tree on every merge since. Two of the 28 — K0 and K1 — predate the one-PR-per-package rule, and two — E1 and X1 — are measurements whose evidence is a document; all four are named in the list above.

  • FIX-590
    Replay is a function of the ledger, not of sibling timing (A1 at concurrency > 1)

    Found by the second POC gate and confirmed by the independent audit: at concurrency above one, the A1 check resolved a claim’s cited URL against any task.retrieved row of the run, so a task’s admission depended on which sibling tasks had already committed — and the same output was rejected “A1: url not fetched in run” when the run was replayed at concurrency 1. The fix: A1 resolves against the task’s own task.retrieved rows (its fetches and its recorded memo hits); a URL the task never retrieved is “not fetched” regardless of siblings — replay-stable by construction, the same shape as input_hash and the memo hit. The raw output of a rejected anchor is kept as evidence, as today.

    done when over a copy of the gate ledger, check 8’s unequal set is exactly the one live run d7f8f21a…, the other 12 replays id-equal and claim_status-equal at 0 calls and 0 reserves, and check 4 still 0 mismatches; a new offline test at concurrency 6 in which one task cites a page only a sibling fetched gives the same admission outcome as at concurrency 1 and a report-equal replay; the 8 gold-8 replays stay byte-identical; SPEC erratum C-186, DESIGN §12 and the TESTPLAN row written issue #640

seq 4Planstage mvp · 56 packages · 55 in

MVP: 55 of 56 in.

One merge commit per package, each verified by three independent readers before it lands. The gate’s own measurement has since been taken — 53 of 53 questions for $8.27, on the MVP page. Identity moved out of the MVP on 17 September by a G3 decision; the fallback ladder and the engines ledger took its place in the gate. One package is open: ONBOARD-1 (#1214) — a user’s first hour, from the request to the first report, every step an event. Its two dependencies, B-SPEND and CONSOLE-1, are both in.

By lane

  • kernelThe kernel itself: consistency classes and the hybrid clock, schema upcast, rejection and repair, driven mode, provenance columns, the kernel-events crate, the substrate adapter, knowledge as a mergeable ledger, and BLOCK-1, which refuses a blocked source at the fetch lane by work identity.15 packages · all in — K6, K7, K8, K9, S1a, FIX-686, S2a, K13, K14, FIX-ARR, FIX-SPIN, QUAL-REPORT, GOLD-SEAT, BLOCK-1, FIX-1281
  • retrievalReal search and fetch: the substrate, the JS-wall fallback, the knowledge producer, index and ingest, the rung-0 fallback chain with its engines ledger, free quota-aware search, and QUAL-READ, which tiers candidates by host and reads three query-focused windows of a page instead of its first 12,000 characters.12 packages · all in — S1, S2, S2b, S2c, S2d, S3, S6, S7, RETR-1, QUAL-READ, FIX-1163, FIX-1211
  • brainPlanning, depth and writing: knowledge-hit guards, campaign memory, QUAL-A and QUAL-B out of the read of the 53 reports, DEPTH-1’s cover loop, NARR-1’s article that has to cite to exist, and JUDGE-1, one verdict row per rubric item.8 packages · all in — B5, B6, QUAL-A, QUAL-B, FIX-1267, DEPTH-1, NARR-1, JUDGE-1
  • evalHow anything here is allowed to claim a number: the Eval ledger, two judges and a sampled citation audit, the mechanical laws, the promotion gate, the gate harness, and GOLD-EXT, the external reference tasks whose questions and rubrics other people wrote.11 packages · all in — E2, E3, E7, E4, X2, X2-RUN, FIX-965, X2-SAFE, E3-MECH, E7-JUDGE, GOLD-EXT
  • governanceMoney and decisions: the budget coordinator, the Policy and Decision ledgers, the five substrate decisions, and B-SPEND — money outside the process, per-user leased keys with limits and a kill switch.4 packages · all in — G1, G2, G3, B-SPEND
  • frontendsThe ways in: OpenCode, the Claude Code skill, rs research, CONSOLE-1 — ask, the run drawn live, the report openable to its bytes — and ONBOARD-1, the one still open.5 packages · open: ONBOARD-1 — F2, F3, F4, CONSOLE-1, ONBOARD-1
  • opsBackup, service and doctor — one package. The CI and build-speed packages share this lane and are counted apart from the MVP, below.1 package · all in — O1

A separate six packages carry CI and build speed — BUILD-1, one integration-test binary for the kernel in place of 59, is in; the umbrella and four others are not started. They are counted apart from the MVP because none of them changes what the tool does, and finding #1401 — the macOS leg of CI failing on the last three pushes to main in about six seconds with no steps recorded — sits in that lane.

Out of reading the 53 reports

Three readers read all 53 gate reports and wrote down 264 defects in 13 root causes (docs/X2-REVIEW-2026-09-20.md). Four packages came out of that reading, and all four are in.

  • QUAL-A
    The audit and the charter know what an expert would read

    The auditor names the canonical sources — book, standard, maintainer page, paper, index, regulation — and a coverage contract over the question's own entities and dimensions that the charter is held to; every prompt knows the run date; a quantitative question is marked so the synthesist must show its arithmetic.

    measured the reference re-recorded on the new auditor through free search: 8 / 8, $0.99, 0 misses · causes #1, #6, #7, half of #8 issue #1106

  • QUAL-B
    The synthesis is gated by the verdict ledger and does the arithmetic

    An answer may rest only on accepted claims; high confidence only when every load-bearing claim is accepted and observed; every number and name in the answer must be entailed by a cited claim's quote or page; worked arithmetic is recomputed by the kernel; a scout's surprise is anchored or rendered as unverified.

    measured the first live pass refused 3 of 8 reference answers, the rule tightened, 8 / 8 after, $1.08 · causes #4, #5, half of #8 issue #1107

  • QUAL-READ
    The run reads the right pages, and reads them properly

    Candidates are tiered by host — maintainer, standard, regulator, paper first; SEO farms last — before the pick; investigators read up to three query-focused windows per page instead of its first 12,000 characters; PDFs, Markdown and XHTML are read; code spans survive extraction.

    merged 2026-09-21 (#1203) after the reference re-recorded with tiered picks · causes #3, #9, #10 issue #1149

  • QUAL-REPORT
    The verifier's two verdicts reconciled, the report clean, knowledge reuse judged

    A page cannot contest itself; the report skips empty sections and de-duplicates gaps; a knowledge hit must pass a judged relevance gate before a line is resolved from it.

    merged 2026-09-21 · causes #2, #11, #13 issue #1108

Still open from that reading: root cause #12, community and documentation hosts refused at rung 0 with only an archive capture behind them — 40 refused trails in the gate, 10 of them reddit. It waits on a Stack Exchange fetch rung (#1109); the other half of it, a search engine’s silent failure surfaced rather than hidden, closed as FIX-1163.docs/PACKAGES.json stage mvp · docs/X2-REVIEW-2026-09-20.md · docs/GATE-MVP.md — failed fetches (S6) · issues #1109, #1163, #1401

seq 5Planfiled 23 Sep · being built

Next, and why these two.

The Jev measurement of 23 September killed the fix we were about to build and named two others. It found 207 of the clean ten’s 500 information-recall items already sitting in pages the runs had downloaded and stored — 30.8 benchmark points against the 16.42 the tool scores — and then found that no single passage carries a complete fact: real items score 0.48–0.64 when the question is put to one passage and 0.96–0.98 when it is put to the whole page. So the answer is not reading more. It is taking more off the page we already read, and then joining what we took.

what is next · filed 23 September · the order their done-whens force the second needs the first; the third is how either is allowed to claim anything EXTRACT-1 #1404 every checkable statement off a page we already read ASSEMBLE-1 #1405 join claims across pages into rows the task asked for re-score gold-ext/v3 the clean ten, scored again, stated beside 16.42 claims per deep run 24.8 → ≥ 60 212 claims from 85 tasks today p50 715 output tokens of 8,000 folds in finding #1326 a record: one row, every cell cited 44 of drb2-024's 96 items are rows all 23 of drb2-062's are rows needs EXTRACT-1 a fall is a finding, not a failure the funnel says which stage moved neither package closes without it ONBOARD-1 #1214 · in parallel · the one open MVP package · a user's first hour, every step an event needs B-SPEND and CONSOLE-1, both merged why these two: 207 of the clean ten's 500 recall items were already in pages the runs had downloaded — 30.8 benchmark points against the 16.42 scored · docs/GOLD-EXT-CLEAN.md · jev.html
  • EXTRACT-1
    Exhaustive, entity-keyed extraction — take every checkable statement off a page we already read

    The investigator produced 212 claims from 85 tasks using a median of 715 output tokens of the 8,000 it is allowed, on a prompt that asks for findings “in roughly {steps} steps”. The window budget everyone blamed accounts for at most 33 of the missed items, 4.9 points. The prompt and schema change to demand every checkable statement with its verbatim quote, keyed — entity, attribute, value, unit, as_of — and the two ceilings downstream of it move with it: the writer’s $0.15 cap refuses a run above about 185 claims, and the coverage CTE counts only claims from an investigator task. Finding #1326 is folded in: a failed article leaves the run done with its report, not gap.

    done when claims per deep run rise from a mean of 24.8 to at least 60 on the clean ten with every quote resolving byte-exactly; the Jev funnel re-run over the new runs shows the “shown but not extracted” class at least halved, its fabricated-fact control re-run beside it; the clean ten re-scored under gold-ext/v3 with the overall stated beside 16.42 — a fall is a finding, not a failure issue #1404

  • ASSEMBLE-1
    Records, not passages — join claims across pages into rows, and render the tables the task asked for

    Asked whether any single passage carried a complete rubric fact, real items scored 0.48–0.64 where fabricated controls scored 0.04; the same question put to the whole page found them at 0.96–0.98. A study's name sits in one place and its country, design and sample size in another, often on another page. A record kind joins EXTRACT-1's entity-keyed claims across passages into a row, every cell carrying the claim it came from, a cell with no claim left empty and named rather than guessed. Much of the external rubric is this shape: 44 of drb2-024's 96 items are rows of two tables the task named, and all 23 of drb2-062's are rows of one.

    done when a record joins claims from at least two different pages into one row on the clean ten with every cell tracing to its claim; the table-shaped items measured before and after; the deliverable contract the task states is extracted and gated like the coverage contract; the clean ten re-scored beside EXTRACT-1's number and 16.42 issue #1405

Neither is in docs/PACKAGES.json on main yet — they were filed as issues on 23 September, after the measurement that justified them, and PR #1406 adds them to the file. That is why the file still counts 108 packages and this page counts 110. The window budget everyone blamed — the one we were about to raise — accounts for at most 33 of the missed items, 4.9 points; that is the fix the measurement killed.docs/JEV-MEASUREMENT.md — the 207 of 500, the 30.8 points and the 33-item window budget · issues #1404, #1405, #1406 — the packages · docs/GOLD-EXT-CLEAN.md · the Jev measurement (jev.html)

seq 6Notefindings · open

What is known wrong with what is built.

A finding is a defect written down against something already merged. Every package here is read adversarially before it lands and again after, and what the reading turns up that is not high severity is filed rather than fixed on the spot — so the honest total is large and worth printing: on 23 September the repository has 1,114 open issues, 998 of them findings, counted as the issues whose title is [<PACKAGE>-F<n>]. Most are narrow — a surviving mutant, an untested branch, a stale TESTPLAN row — and none of them is a reason to trust a number here less, because each one is a defect this project found in itself and wrote down.

The five below are different: each is visible in a report or in a number published on this site, so a reader of it should know about them by name.gh issue list --state open over CPUtester5465/research-stack, 2026-09-23 · the finding titles are the adversarial reading’s own convention

  • #1326An article the writer cannot get past the gate ends the whole run gap, discarding a good synthesis. It cost three of ten runs in the first external pass and drb2-024 again in the clean one, each with a finished synthesis on the record. Folded into EXTRACT-1: a failed article will leave the run done with its report and the refusal named in its meta line; gap is for research that could not be completed.open · NARR-1-F8
  • #1327The report pads — a fetch-failure log, a Findings list that repeats Sources, and “what is not established” printed twice. Presentation scores 27.04 on the clean ten, and this is part of why.open · NARR-1-F9
  • #1328Analysis halves when the cover loop doubles the claims: DEPTH-1 took retrieval targets from 91.5 to 191.7 and claims per task from 14.6 to 24.8, and information recall rose 99 % while analysis fell 47 %. The synthesist’s budget does not scale with what it is given. Plan v1 is still the default because of it.open · DEPTH-1-F1 · docs/DEPTH-1-NOTES.md
  • #1401The macOS leg of the CI workflow has failed on the last three pushes to main, about six seconds each with no steps recorded. It does not touch any number on this site — the gate checks that run in CI are the Linux leg — but it is red and it is ours.open · CI
  • #1330Four of ten gold-ext runs read the source their own task forbids. Fixed and measured: BLOCK-1 merged on 22 September (#1399) and enforces the blocklist at the fetch lane by work identity; in the 23 September re-run it refused three fetches, one per mechanism it was built for — the /pdf/ twin of an id listed as /abs/, and two listed URLs. The four tasks were re-run without the forbidden source and are the four clean rows behind the 16.42 this site publishes.the defect is closed in the code; the issue is still open

The four re-run tasks scored higher without the source their rubric came from — 16.76 with it, 17.61 without. Reading the answer key was not an advantage, which is worth remembering the next time a shortcut looks like one.issues #1326, #1327, #1328, #1330, #1401 · docs/GOLD-EXT-CLEAN.md · docs/DEPTH-1-NOTES.md

seq 7Planstage prod · 18 packages

Then production.

None of the 18 has started, and none of them is a condition of the MVP being measured — they are what stands between a measured MVP and something other people can run. Four of them are the ones a reader asks about.

  • K10
    Multi-node sync

    The mergeable class across the Mac and a cloud node; the sequenced class log-shipped from its owner; either side may be offline.

    done when a replica rebuilt from the other node yields identical derived-view hashes after 24 h of divergent writes issue #37

  • K11
    Transparency log

    A Merkle tree over events as a view; hourly checkpoints anchored for $0; rs seen returns an inclusion proof.

    done when a proof verifies offline against a published checkpoint issue #38

  • K12
    Tiering to R2

    Primary ledgers age to Parquet on R2 with zone maps; derived ledgers drop and rebuild.

    done when a year-old run replays from the cold tier with the same hash issue #39

  • F5
    Cloud clients

    Streamable-HTTP MCP with bearer auth against the cloud node; per-client fairness.

    done when a driven run from a remote Claude Code session completes against the cloud node issue #44

The rest, by title: B7 Plan v2: skippable phases · B8 Graph retriever · S4 GLiNER extraction · S5 Change-cadence verification · E5 Off-policy eval and multileaving · E6 Red-team closure · G4 Learned budgets · G5 Expiry and PROV-O export · O2 Cluster deploy · O3 SLOs and alerts · O4 Docs and licence review · X3 Production gate measurement · I1 rs-identity — identity orchestrator + vault (rung 1) · I2 rs-browser — CDP worker, home-node-pinned. I1 rs-identity and I2 rs-browser are here rather than in the MVP because a G3 decision moved them on 17 September.docs/PACKAGES.json, stage prod · docs/PACKAGES.json decisions, 2026-09-17

next: try it — who it is for, and how the MVP is driven today.