← ALL POSTS · /BLOG

RELEASE GATES · TEN FEATURES · EIGHT GATES · RUNS 2026-09-30

THE RC.80 GATES.
MEASURED IN PUBLIC, LOSSES LEFT IN.

THE VERDICT split, measured, with the losses left in

The program behind rc.80 shipped ten features, and every item that could carry a number was filed with a measured gate that binds at ship time. This post is the packaging half of that rule, and the scoreboard is a split. Three gates passed outright: /_ask turned 230 prompt/gold pairs into validated DSL at macro-F1 0.9975 with zero invalid queries and wall-clock p50 1.129 ms against a 300 ms budget; xerj_map took unknown-field 400s on MCP-shaped DSL from 20/30 guessed to 0/30 map-informed — a 100% reduction against a ≥80% bar; and the calibration layer took FiQA rerank ECE from 0.3109 raw to 0.0088 held-out against a ≤0.10 bar. Three failed with the numbers: the local judge loses to its own first stage everywhere it was measured (SciFact 0.7045 → 0.5967), the decision flywheel answers 19.72% of later traffic at ≥0.8 confidence against an 80% bar, and the tier-2 decide head misses its accuracy bars by 3 to 59 points. One split: stemming gained +0.0160/+0.0179 BM25 nDCG@10 (bar +0.01) and still left 15 zero-hit queries against a bar of ≤10 — and no stemmer can fix the 15. The failures filed four issues, #1091–#1094. No number below was tuned after its gate was read.

One window, ten features, eight gates, all measured on 2026-09-30 against bars that were written down before the work started. The wins and the losses sit in the same tables, every figure links to the run that produced it, and the four defects the gates exposed are filed in the open tracker with the measurements that found them. That is the whole post.

0.9975
macro-F1, /_ask
230 prompt/gold pairs, 0 invalid DSL, byte-identical across 3 runs. p50 1.129 ms; the bar allowed 300.
0/30
unknown-field 400s, with xerj_map
down from 20/30 guessed field names. 100% reduction; the bar asked for 80%. The 5 silent 0-hits became issue #1093.
0.0088
held-out ECE, calibrated
from 0.3109 raw on 3,600 held-out FiQA pairs, query-level split. Temperature scaling managed 0.3163 and is published beside it.
4
issues the failures filed
#1091 hybrid cost, #1092 analyzer echo, #1093 silent 0-hit, #1094 flywheel freeze. Each carries its measurement.
01·THE PROGRAM · TEN FEATURES, EACH WITH A GATE THAT BINDS

THE RULE WAS WRITTEN
BEFORE THE WORK.

The ten-item program to the first official release (discussion #1054, filed as issues #1055–#1064) was planned as five items for rc.79, four for rc.80, and the open model for v1.1.0. They landed together in one window — PRs #1066–#1090 since the rc.78 cut — so the gates were run as one suite, and the artifacts call it the rc.80 gate suite. This post keeps that name and states the plan honestly: two planned releases, one measured window.

The plan's own packaging rule, verbatim: a benchmarks/<name> directory with raw results plus a blog post with the losses left in. Three items also carried a loses-does-not-ship rule in their issue text — the local judge must beat the shipped hybrid arm beyond the run spread or it does not ship; the decide head below its bar ships behind a flag with the numbers published; the FiQA judge row must clear 0.30. This post is where those rules get honoured.

The program · feature (PR) · its gateVerdict
xerj_map MCP FieldSpec tool (#1074) · unknown-field 400s down ≥80% PASS — 100%
POST /_ask + xerj_plan (#1076) · macro-F1 ≥0.9, 0 invalid DSL, p50 ≤300 ms PASS
tier-2 local decide head via candle (#1072, #1073) · SMS ≥0.95 acc, ECE ≤0.05, AG News ≥0.85 FAIL (quality) / PASS (latency)
max_tokens on every MCP search tool (#1067) · response-size budget knob no measured gate in this suite
stemming analyzer for profiler-marked prose (#1070) · BM25 ≥+0.01; zero-hit ≤10 SPLIT — +0.0160/+0.0179 PASS; 15 zero-hit FAIL
self-judged hits judge stage (#1077) · beat hybrid beyond spread, or do not ship FAIL (quality) / PASS (cost)
decision-cache flywheel (#1075) · ≥80% of later traffic from history at ≥0.8, acc ≥0.97 FAIL (share) / PASS (accuracy)
real _watcher evaluate-or-501 + autoindex --label (#1082) · wrong-and-confident ≤5% at p≥0.8 gate not yet measured
calibration layer, p_cal (#1080) · held-out ECE ≤0.10 PASS — 0.0088
open xerj-decide-v1 model (#1084) · deterministic exports, published eval card reported — card carries the weak numbers

The plan: discussion #1054 · tracker #1065 · issues #1055–#1064 · window PRs #1066–#1090 · gate PRs #1079, #1085, #1086, #1087, #1088, #1090.

02·THE PASSES · THREE GATES, CLEAN

THREE PASSED.
WITH THEIR NUMBERS.

PASS 1

POST /_ask + xerj_plan: prompt in, validated DSL out

#1056 · PRs #1076, #1079

The gate: ≥200 (prompt, gold result-set) pairs over public tabular datasets, result-set F1 ≥0.9, zero invalid DSL, and a latency budget. Measured: 230 pairs (72 USGS earthquakes / 78 exoplanets / 80 gapminder), macro-F1 0.9975, 0 invalid DSL — every returned plan passes parse_request before it is returned — and byte-identical output across 3 runs. Wall-clock p50 1.129 ms, p99 2.852 ms, n=690 over loopback HTTP. The bar allowed 300 ms; the planner clears it by 265×.

The one arm not run: the token gate — output tokens per solved structured-query task ≤50% of agent-written DSL at equal solve rate — is designed, not measured. It needs a live agent harness with real model calls, and the cost of that run was not spent this window. The design sits in benchmarks/ask-plan and the claim is not made here.

PASS 2

xerj_map: the field map an agent was guessing at

#1055 · PRs #1074, #1086

Thirty MCP-shaped DSL queries against a reference-coding corpus, issued twice: once guessing field names from the prompt alone, once informed by the map. Guessed: 20 of 30 came back as unknown-field 400s, and only 5 returned hits. Map-informed: 0 of 30 unknown-field 400s, 29 of 30 with hits. A 100% reduction against the ≥80% bar.

The discovery is worth more than the pass. Five of the guessed queries did not error at all — a wrong field name inside term, match, range or exists is a silent 0-hit HTTP 200. The gate counted 400s; the silent rows are the worse failure, and they are now issue #1093.

PASS 3

Calibration: p_cal beside every p_raw

#1063 · PRs #1080, #1087

The baseline is the number this blog already published: on FiQA, a hosted judge probability of 0.93 meant relevance 34% of the time — ECE 0.3109. The gate: held-out ECE ≤0.10 after fitting. Measured on 19,440 raw (query, doc, p_raw, gold) pairs rebuilt from 648 judged FiQA queries — $0.2125 of provider calls, 5,059,586 input tokens, zero failed calls — with a query-level split: the fit never sees any pair from a held-out query, which is the deployment shape. Isotonic held-out ECE 0.0088. Temperature scaling, the comparison arm, managed 0.3163 and fails the same bar; both arms' ECE is published by the node at /_decide/_calibration.

The gate also audited the fitter. Writing an independent naive PAVA in Python found a knot-index bug in the benchmark's mirror — fixed there, and the shipped Rust module was verified against the naive implementation to 8.7e-19 at all 98 pooled points. The defect was in the measuring stick, not the product; both are now pinned.

ask-plan: benchmarks/ask-plan (gate table with the measured rows) and results/2026-09-30-latency-gate · map gate: PR #1086 (artifacts in the PR) · calibration: results/2026-09-30-pairlevel-fiqa · pairs and shortlists committed beside them.

03·THE FAILURES · READINGS, NOT EXCUSES

THREE FAILED, ONE SPLIT.
THE NUMBERS STAY IN.

Stemming (#1059): a real win and a real miss, in the same gate

Paired arms on one node, differing only in the create-time analyzer. BM25 nDCG@10: SciFact 0.6572 → 0.6732 (+0.0160), NFCorpus 0.3016 → 0.3195 (+0.0179) — both clear the +0.01 bar. Zero-hit NFCorpus queries: 25 → 15, against a bar of ≤10. Stemming repaired exactly the ten plural and inflection queries (bagels, leeks, pineapples, turnips…). The fifteen that remain — Fosamax, Zoloft, eggnog, halibut, Yale — are single words whose stem also occurs nowhere in the corpus. No stemmer can lexically bridge a word that is absent in every inflected form. That half of the gate was unreachable by the feature it gated; only the vector arm can fix those queries, and that is the honest reading: the gate was half right to exist, and the miss is a boundary, not a bug in the stemmer.

Two adjacent records: GET _settings does not echo the analysis block and _analyze shows the standard path even when the stemmer is declared and provably applied by search — filed as issue #1092 rather than tuned around. Raw per-query outputs are committed.

The tier-2 decide head (#1057): misses its bars, ships behind the flag anyway

Decide head, xerj-decide-v1 weights · measured on the wireAccuracyECEBarVerdict
SMS held-out, noul · 1,574 items0.91930.287acc ≥0.95, ECE ≤0.05FAIL / FAIL
AG News test, 4-way choice · 7,600 items, never trained on it0.25990.060≥0.85FAIL (chance = 0.25)
Banking77 test, 77-way choice · 3,080 items0.11850.0717reported, not gatedreported
Latency, 30 questions, 8 pinned coresp50 8.4 ms≤50 msPASS

The reading is in the issue's own rule: below the bar it ships behind a flag with the numbers published — --decide-mode local --decide-model-dir, tier 2 off unless asked for. The artifact is 28.5 MiB of F32 weights, not quantized (candle 0.9 has no quantized safetensors VarBuilder; the PR documents that as deliberate), well under the 300 MB download budget — and there is no download at all: the tier loads a local directory and makes no egress. Tier 1, the history vote, stays ahead of tier 2 on every home-ground row: Banking77 0.8289 vs 0.1185, SMS 0.9848 vs 0.9193, AG News 0.9182 vs 0.2599. That is the design working: where labelled history exists, the vote wins; where it does not, tier 2 exists at all — and its zero-shot boundary is now a published number instead of a hope.

The gate's second catch is the more useful one. Measuring tier 2 on the stock release binary found that --decide-mode local was a no-op there: release binaries were built without the decide-local cargo feature, the node logged a warning, and the ladder quietly stayed on the history vote. Every tier-2 number above was therefore measured on a feature build of the same commit, with the no-op evidence committed (stock-binary-noop.txt). The fix is release engineering, not code: the shipped build must carry the feature — the runtime switch stays opt-in — and it is queued ahead of the cut. A gate that cannot reach the feature it is gating is measuring the build, and that is exactly what it caught.

The flywheel (#1061): frozen as shipped, 4× under the bar even when forced

Two arms, both worse than the bar. As shipped: 2,000 fill requests fired, and only 12 answers were ever cached — write-back fires only when tier 2 answers, and after the first cached document the history vote returns BM25 support for ~99.4% of traffic, so tier 2 stops answering and the cache freezes at 12 documents. That is a defect, filed as issue #1094, and it makes the gate's premise unreachable through the ladder itself. The premise forced — 2,000 gold answers cached in the write-back's own document shape — later traffic answered from history at ≥0.8 confidence: 19.72% against the 80% bar, four times under, at accuracy 0.9953 against the 0.97 bar, which met. The reading, not the excuse: the confident slice is easily accurate enough, but 2,000 answers at k=10 only clear 0.8 confidence on a fifth of Banking77-shaped traffic. The share bar needs a far larger cache, or the hybrid neighbours, before it is reachable — and until #1094 is fixed the cache cannot grow at all.

The local judge (#1060): loses to its own first stage, everywhere measured

nDCG@10 · 3 runs each, spread 0.0000First stage+ judgeΔBarVerdict
SciFact · 300 q0.7045 hybrid0.5967−0.1077beat hybrid beyond spreadFAIL
NFCorpus · 323 q0.3419 hybrid0.2924−0.0495beat hybrid beyond spreadFAIL
FiQA · 648 q0.2382 BM250.1650−0.0732≥0.30 (hosted Jev: 0.3638)FAIL
Added latency, p50, top-30+0.7 ms / +0.5 ms / +0.6 ms≤40 msPASS

Why it loses, plainly: the shipped judge is a lexical scorer — window-local BM25 saturation with page-local IDF, no semantic signal. Both first stages already read those words: BM25 ranked by them, hybrid fused them with MiniLM vectors. Re-reading the same words over the page with page-local statistics destroys information the first stages had — corpus-level IDF and the vector arm — so the judged order is a worse order. The effect is biggest exactly where the first stage is strongest. The gate arms carried no min_p, because a threshold can only remove recall and cannot be part of a passing attempt; an informational arm with min_p: 0.5 kept 0.1 of 30 hits per page on SciFact and emptied 274 of 300 queries. The stage stays what PR #1077 shipped: opt-in, named lexical in every response, no quality claim — its own loses-does-not-ship rule, honoured. One cell could not be measured at all: hybrid+judge on FiQA, because the hybrid first stage runs 13–26 s per query on the 57,638-doc index — issue #1091, found by this gate.

Stemming: results/2026-09-30-stemming-1059 (PR #1085) · decide + flywheel: PR #1088 — benchmarks/decisions-as-retrieval/results/2026-09-30-rc80gates/, incl. run-2026-09-30.md and stock-binary-noop.txt · judge: PR #1090 — benchmarks/beir-hybrid/results/2026-09-30-judge-gate/, incl. the aborted FiQA hybrid log behind #1091.

04·THE NEXT WINDOW · WHAT rc.81 DOES ABOUT EACH LOSS

EVERY LOSS HAS
A FILED ISSUE.

A failed gate that files nothing is a buried gate. These filed four, plus two direction changes the numbers forced:

The lossThe filed work
silent 0-hit on unknown fields (found by the map gate)#1093 — QueryError::UnknownField is constructed nowhere today; wrong fields must 4xx, not 200-empty
flywheel freeze (12 of 2,000 cached)#1094 — fix the write-back rule so the cache can grow; the 19.72% share measurement is then re-run on a live cache
hybrid 13–26 s/query on FiQA (blocked the judge's FiQA hybrid cell)#1091 — first-stage cost on the 57,638-doc index; the hosted-comparison row needs it
analyzer invisible in settings/_analyze#1092 — echo the declared analysis block
lexical judge loses everywherethe model arm: a judge that ships must be the decide-local cross-encoder head, beating these same hybrid numbers under this same harness — the bar does not move
the 15 unstemmable zero-hit queriesthe vector arm — the only repair for words absent from the corpus in every inflected form

Two more states, stated so nobody has to guess. The open model: xerj-decide-v1 is 7,419,907 parameters, its exports are byte-identical across runs, its eval card carries the weak numbers above in it, and the Hugging Face publication is staged behind an operator credential — no URL is claimed until the page exists. And the decide-head build fix: the shipped binary must carry the decide-local feature so the opt-in flag is not a no-op; that change rides ahead of the cut.

What we claim, and what we do not. Claimed: every number on this page, verbatim from the gate runs linked beside it, on release builds of main @ a179e1ef3 measured 2026-09-30 — the tier-2 numbers on the feature build of that commit, the calibration pairs on one paid provider run whose determinism bound is stated in its provenance file. Not claimed: any quality for the judge stage, any quality for tier 2 beyond the published numbers, any token-saving result for xerj_plan (designed, not run), any detection-quality figure for _watcher (gate not yet measured), and any published URL for the model. The gates were written before the work; three said yes, three said no, one split; the repo carries all eight verdicts and the four issues the noes filed.

If you want to check a number rather than read about it: each gate directory above carries its raw per-query outputs, its manifest and its harness, and the four gate PRs (#1086, #1088, #1090, and #1085's merged artifacts) are diffable line by line. The engine itself is one binary with a quickstart, the playground needs no install, and the benchmarks hub keeps the standing boards. The nearest prior chapter is the Jev verdict — the 0.3109 ECE this window's calibration gate finally answered.

Everything on this page: the gate runs cited section by section above · program #1054 · gate PRs #1079 #1085 #1086 #1087 #1088 #1090 · issues #1091–#1094 · related: the Jev verdict · the benchmarks hub.