THE RC.80 GATES.
MEASURED IN PUBLIC, LOSSES LEFT IN.
The program behind rc.80 shipped ten features, and every item that could carry a number was filed with a measured gate that binds at ship time. This post is the packaging half of that rule, and the scoreboard is a split. Three gates passed outright: /_ask turned 230 prompt/gold pairs into validated DSL at macro-F1 0.9975 with zero invalid queries and wall-clock p50 1.129 ms against a 300 ms budget; xerj_map took unknown-field 400s on MCP-shaped DSL from 20/30 guessed to 0/30 map-informed — a 100% reduction against a ≥80% bar; and the calibration layer took FiQA rerank ECE from 0.3109 raw to 0.0088 held-out against a ≤0.10 bar. Three failed with the numbers: the local judge loses to its own first stage everywhere it was measured (SciFact 0.7045 → 0.5967), the decision flywheel answers 19.72% of later traffic at ≥0.8 confidence against an 80% bar, and the tier-2 decide head misses its accuracy bars by 3 to 59 points. One split: stemming gained +0.0160/+0.0179 BM25 nDCG@10 (bar +0.01) and still left 15 zero-hit queries against a bar of ≤10 — and no stemmer can fix the 15. The failures filed four issues, #1091–#1094. No number below was tuned after its gate was read.
One window, ten features, eight gates, all measured on 2026-09-30 against bars that were written down before the work started. The wins and the losses sit in the same tables, every figure links to the run that produced it, and the four defects the gates exposed are filed in the open tracker with the measurements that found them. That is the whole post.
THE RULE WAS WRITTEN
BEFORE THE WORK.
The ten-item program to the first official release (discussion #1054, filed as issues #1055–#1064) was planned as five items for rc.79, four for rc.80, and the open model for v1.1.0. They landed together in one window — PRs #1066–#1090 since the rc.78 cut — so the gates were run as one suite, and the artifacts call it the rc.80 gate suite. This post keeps that name and states the plan honestly: two planned releases, one measured window.
The plan's own packaging rule, verbatim: a benchmarks/<name> directory with raw results plus a blog post with the losses left in. Three items also carried a loses-does-not-ship rule in their issue text — the local judge must beat the shipped hybrid arm beyond the run spread or it does not ship; the decide head below its bar ships behind a flag with the numbers published; the FiQA judge row must clear 0.30. This post is where those rules get honoured.
| The program · feature (PR) · its gate | Verdict |
|---|---|
xerj_map MCP FieldSpec tool (#1074) · unknown-field 400s down ≥80% |
PASS — 100% |
POST /_ask + xerj_plan (#1076) · macro-F1 ≥0.9, 0 invalid DSL, p50 ≤300 ms |
PASS |
| tier-2 local decide head via candle (#1072, #1073) · SMS ≥0.95 acc, ECE ≤0.05, AG News ≥0.85 | FAIL (quality) / PASS (latency) |
max_tokens on every MCP search tool (#1067) · response-size budget knob |
no measured gate in this suite |
| stemming analyzer for profiler-marked prose (#1070) · BM25 ≥+0.01; zero-hit ≤10 | SPLIT — +0.0160/+0.0179 PASS; 15 zero-hit FAIL |
| self-judged hits judge stage (#1077) · beat hybrid beyond spread, or do not ship | FAIL (quality) / PASS (cost) |
| decision-cache flywheel (#1075) · ≥80% of later traffic from history at ≥0.8, acc ≥0.97 | FAIL (share) / PASS (accuracy) |
real _watcher evaluate-or-501 + autoindex --label (#1082) · wrong-and-confident ≤5% at p≥0.8 |
gate not yet measured |
calibration layer, p_cal (#1080) · held-out ECE ≤0.10 |
PASS — 0.0088 |
| open xerj-decide-v1 model (#1084) · deterministic exports, published eval card | reported — card carries the weak numbers |
- Two rows say "not measured" on purpose.
max_tokensis a response-size budget with no quality claim attached, and the_watcherdetection gate (wrong-and-confident rate at p≥0.8) has not been run — the watcher's mechanics are test-pinned, its detection-quality number is not, and this post does not invent one. - One binary, one day. Every number in this post was measured on 2026-09-30 on release builds of main @
a179e1ef3, version stringv1.0.0-rc.78— the rc.79/rc.80 tags are cut after the gates, not before them. The two exceptions are stated where they occur (the tier-2 feature build, the paid provider arm).
The plan: discussion #1054 · tracker #1065 · issues #1055–#1064 · window PRs #1066–#1090 · gate PRs #1079, #1085, #1086, #1087, #1088, #1090.
THREE PASSED.
WITH THEIR NUMBERS.
POST /_ask + xerj_plan: prompt in, validated DSL out
#1056 · PRs #1076, #1079The gate: ≥200 (prompt, gold result-set) pairs over public tabular datasets, result-set F1 ≥0.9, zero invalid DSL, and a latency budget. Measured: 230 pairs (72 USGS earthquakes / 78 exoplanets / 80 gapminder), macro-F1 0.9975, 0 invalid DSL — every returned plan passes parse_request before it is returned — and byte-identical output across 3 runs. Wall-clock p50 1.129 ms, p99 2.852 ms, n=690 over loopback HTTP. The bar allowed 300 ms; the planner clears it by 265×.
The one arm not run: the token gate — output tokens per solved structured-query task ≤50% of agent-written DSL at equal solve rate — is designed, not measured. It needs a live agent harness with real model calls, and the cost of that run was not spent this window. The design sits in benchmarks/ask-plan and the claim is not made here.
xerj_map: the field map an agent was guessing at
#1055 · PRs #1074, #1086Thirty MCP-shaped DSL queries against a reference-coding corpus, issued twice: once guessing field names from the prompt alone, once informed by the map. Guessed: 20 of 30 came back as unknown-field 400s, and only 5 returned hits. Map-informed: 0 of 30 unknown-field 400s, 29 of 30 with hits. A 100% reduction against the ≥80% bar.
The discovery is worth more than the pass. Five of the guessed queries did not error at all — a wrong field name inside term, match, range or exists is a silent 0-hit HTTP 200. The gate counted 400s; the silent rows are the worse failure, and they are now issue #1093.
Calibration: p_cal beside every p_raw
#1063 · PRs #1080, #1087The baseline is the number this blog already published: on FiQA, a hosted judge probability of 0.93 meant relevance 34% of the time — ECE 0.3109. The gate: held-out ECE ≤0.10 after fitting. Measured on 19,440 raw (query, doc, p_raw, gold) pairs rebuilt from 648 judged FiQA queries — $0.2125 of provider calls, 5,059,586 input tokens, zero failed calls — with a query-level split: the fit never sees any pair from a held-out query, which is the deployment shape. Isotonic held-out ECE 0.0088. Temperature scaling, the comparison arm, managed 0.3163 and fails the same bar; both arms' ECE is published by the node at /_decide/_calibration.
The gate also audited the fitter. Writing an independent naive PAVA in Python found a knot-index bug in the benchmark's mirror — fixed there, and the shipped Rust module was verified against the naive implementation to 8.7e-19 at all 98 pooled points. The defect was in the measuring stick, not the product; both are now pinned.
ask-plan: benchmarks/ask-plan (gate table with the measured rows) and results/2026-09-30-latency-gate · map gate: PR #1086 (artifacts in the PR) · calibration: results/2026-09-30-pairlevel-fiqa · pairs and shortlists committed beside them.
THREE FAILED, ONE SPLIT.
THE NUMBERS STAY IN.
Stemming (#1059): a real win and a real miss, in the same gate
Paired arms on one node, differing only in the create-time analyzer. BM25 nDCG@10: SciFact 0.6572 → 0.6732 (+0.0160), NFCorpus 0.3016 → 0.3195 (+0.0179) — both clear the +0.01 bar. Zero-hit NFCorpus queries: 25 → 15, against a bar of ≤10. Stemming repaired exactly the ten plural and inflection queries (bagels, leeks, pineapples, turnips…). The fifteen that remain — Fosamax, Zoloft, eggnog, halibut, Yale — are single words whose stem also occurs nowhere in the corpus. No stemmer can lexically bridge a word that is absent in every inflected form. That half of the gate was unreachable by the feature it gated; only the vector arm can fix those queries, and that is the honest reading: the gate was half right to exist, and the miss is a boundary, not a bug in the stemmer.
Two adjacent records: GET _settings does not echo the analysis block and _analyze shows the standard path even when the stemmer is declared and provably applied by search — filed as issue #1092 rather than tuned around. Raw per-query outputs are committed.
The tier-2 decide head (#1057): misses its bars, ships behind the flag anyway
| Decide head, xerj-decide-v1 weights · measured on the wire | Accuracy | ECE | Bar | Verdict |
|---|---|---|---|---|
| SMS held-out, noul · 1,574 items | 0.9193 | 0.287 | acc ≥0.95, ECE ≤0.05 | FAIL / FAIL |
| AG News test, 4-way choice · 7,600 items, never trained on it | 0.2599 | 0.060 | ≥0.85 | FAIL (chance = 0.25) |
| Banking77 test, 77-way choice · 3,080 items | 0.1185 | 0.0717 | reported, not gated | reported |
| Latency, 30 questions, 8 pinned cores | p50 8.4 ms | ≤50 ms | PASS | |
The reading is in the issue's own rule: below the bar it ships behind a flag with the numbers published — --decide-mode local --decide-model-dir, tier 2 off unless asked for. The artifact is 28.5 MiB of F32 weights, not quantized (candle 0.9 has no quantized safetensors VarBuilder; the PR documents that as deliberate), well under the 300 MB download budget — and there is no download at all: the tier loads a local directory and makes no egress. Tier 1, the history vote, stays ahead of tier 2 on every home-ground row: Banking77 0.8289 vs 0.1185, SMS 0.9848 vs 0.9193, AG News 0.9182 vs 0.2599. That is the design working: where labelled history exists, the vote wins; where it does not, tier 2 exists at all — and its zero-shot boundary is now a published number instead of a hope.
--decide-mode local was a no-op there: release binaries were built without the decide-local cargo feature, the node logged a warning, and the ladder quietly stayed on the history vote. Every tier-2 number above was therefore measured on a feature build of the same commit, with the no-op evidence committed (stock-binary-noop.txt). The fix is release engineering, not code: the shipped build must carry the feature — the runtime switch stays opt-in — and it is queued ahead of the cut. A gate that cannot reach the feature it is gating is measuring the build, and that is exactly what it caught.
The flywheel (#1061): frozen as shipped, 4× under the bar even when forced
Two arms, both worse than the bar. As shipped: 2,000 fill requests fired, and only 12 answers were ever cached — write-back fires only when tier 2 answers, and after the first cached document the history vote returns BM25 support for ~99.4% of traffic, so tier 2 stops answering and the cache freezes at 12 documents. That is a defect, filed as issue #1094, and it makes the gate's premise unreachable through the ladder itself. The premise forced — 2,000 gold answers cached in the write-back's own document shape — later traffic answered from history at ≥0.8 confidence: 19.72% against the 80% bar, four times under, at accuracy 0.9953 against the 0.97 bar, which met. The reading, not the excuse: the confident slice is easily accurate enough, but 2,000 answers at k=10 only clear 0.8 confidence on a fifth of Banking77-shaped traffic. The share bar needs a far larger cache, or the hybrid neighbours, before it is reachable — and until #1094 is fixed the cache cannot grow at all.
The local judge (#1060): loses to its own first stage, everywhere measured
| nDCG@10 · 3 runs each, spread 0.0000 | First stage | + judge | Δ | Bar | Verdict |
|---|---|---|---|---|---|
| SciFact · 300 q | 0.7045 hybrid | 0.5967 | −0.1077 | beat hybrid beyond spread | FAIL |
| NFCorpus · 323 q | 0.3419 hybrid | 0.2924 | −0.0495 | beat hybrid beyond spread | FAIL |
| FiQA · 648 q | 0.2382 BM25 | 0.1650 | −0.0732 | ≥0.30 (hosted Jev: 0.3638) | FAIL |
| Added latency, p50, top-30 | +0.7 ms / +0.5 ms / +0.6 ms | ≤40 ms | PASS | ||
Why it loses, plainly: the shipped judge is a lexical scorer — window-local BM25 saturation with page-local IDF, no semantic signal. Both first stages already read those words: BM25 ranked by them, hybrid fused them with MiniLM vectors. Re-reading the same words over the page with page-local statistics destroys information the first stages had — corpus-level IDF and the vector arm — so the judged order is a worse order. The effect is biggest exactly where the first stage is strongest. The gate arms carried no min_p, because a threshold can only remove recall and cannot be part of a passing attempt; an informational arm with min_p: 0.5 kept 0.1 of 30 hits per page on SciFact and emptied 274 of 300 queries. The stage stays what PR #1077 shipped: opt-in, named lexical in every response, no quality claim — its own loses-does-not-ship rule, honoured. One cell could not be measured at all: hybrid+judge on FiQA, because the hybrid first stage runs 13–26 s per query on the 57,638-doc index — issue #1091, found by this gate.
Stemming: results/2026-09-30-stemming-1059 (PR #1085) ·
decide + flywheel: PR #1088 —
benchmarks/decisions-as-retrieval/results/2026-09-30-rc80gates/, incl. run-2026-09-30.md and stock-binary-noop.txt ·
judge: PR #1090 —
benchmarks/beir-hybrid/results/2026-09-30-judge-gate/, incl. the aborted FiQA hybrid log behind #1091.
EVERY LOSS HAS
A FILED ISSUE.
A failed gate that files nothing is a buried gate. These filed four, plus two direction changes the numbers forced:
| The loss | The filed work |
|---|---|
| silent 0-hit on unknown fields (found by the map gate) | #1093 — QueryError::UnknownField is constructed nowhere today; wrong fields must 4xx, not 200-empty |
| flywheel freeze (12 of 2,000 cached) | #1094 — fix the write-back rule so the cache can grow; the 19.72% share measurement is then re-run on a live cache |
| hybrid 13–26 s/query on FiQA (blocked the judge's FiQA hybrid cell) | #1091 — first-stage cost on the 57,638-doc index; the hosted-comparison row needs it |
| analyzer invisible in settings/_analyze | #1092 — echo the declared analysis block |
| lexical judge loses everywhere | the model arm: a judge that ships must be the decide-local cross-encoder head, beating these same hybrid numbers under this same harness — the bar does not move |
| the 15 unstemmable zero-hit queries | the vector arm — the only repair for words absent from the corpus in every inflected form |
Two more states, stated so nobody has to guess. The open model: xerj-decide-v1 is 7,419,907 parameters, its exports are byte-identical across runs, its eval card carries the weak numbers above in it, and the Hugging Face publication is staged behind an operator credential — no URL is claimed until the page exists. And the decide-head build fix: the shipped binary must carry the decide-local feature so the opt-in flag is not a no-op; that change rides ahead of the cut.
a179e1ef3 measured 2026-09-30 — the tier-2 numbers on the feature build of that commit, the calibration pairs on one paid provider run whose determinism bound is stated in its provenance file. Not claimed: any quality for the judge stage, any quality for tier 2 beyond the published numbers, any token-saving result for xerj_plan (designed, not run), any detection-quality figure for _watcher (gate not yet measured), and any published URL for the model. The gates were written before the work; three said yes, three said no, one split; the repo carries all eight verdicts and the four issues the noes filed.
If you want to check a number rather than read about it: each gate directory above carries its raw per-query outputs, its manifest and its harness, and the four gate PRs (#1086, #1088, #1090, and #1085's merged artifacts) are diffable line by line. The engine itself is one binary with a quickstart, the playground needs no install, and the benchmarks hub keeps the standing boards. The nearest prior chapter is the Jev verdict — the 0.3109 ECE this window's calibration gate finally answered.
Everything on this page: the gate runs cited section by section above · program #1054 · gate PRs #1079 #1085 #1086 #1087 #1088 #1090 · issues #1091–#1094 · related: the Jev verdict · the benchmarks hub.