Five things a search engine for AI agents has to prove, each measured, each linked to the
harness or the case study that produced it: engine performance,
code retrieval, ranking quality, token savings, and the design that ties them
together — the zero-token architecture. The red cells are published too, root-caused with
file:line, because a benchmark page that hides its losses is marketing.
Two performance stories matter for an engine that lives next to your agents: how fast a query comes back, and what the process costs when nobody is asking. Both are measured; neither is extrapolated.
| Measurement | XERJ | Comparison | Verdict | Source |
|---|---|---|---|---|
| 88-cell query board vs Elasticsearch 8.x — reads, aggs, pipelines, mixed, kNN, storage | 55 wins | 26 ties · 4 losses · 3 n/a | WIN | scorecard — one win is honestly a draw; the page says which |
Server round-trip, size:0 over 300k docs, keep-alive client |
0.126 ms | ES 0.283 ms | 2.2× | scorecard · read transport |
| Bulk ingest, identical corpus and box | 1.72× higher | ES baseline | WIN | scorecard |
| On-disk size, same docs | 176.2 MB | ES 283.0 MB | 1.61× smaller | scorecard · storage caveats apply |
| kNN k=10, HNSW-served with exact rescoring | 1.76 ms | ES 2.08 ms | TIE (1.18×, inside noise) | scorecard — recall is measured, never assumed |
| Identifier lookup, one index, 15-repo code corpus (p50 of 21) | 11.7 ms | Zoekt (Sourcegraph) 15.3 ms · ripgrep cold 260 ms | WIN per-index | 2026-08-29 session, same box, released rc.70 binary |
| Mixed read-under-write p99, 40k docs/s iso-load writer | 10.3–13.6 ms | ES 3.5–6.8 ms | 4 LOSSES | scorecard · root-caused to a lock, fix targeted |
| Dimension | XERJ rc.70/71 | Zoekt (Sourcegraph) | ripgrep | What it means for you |
|---|---|---|---|---|
| Identifier search, p50 | 11.7 ms (one index) | 15.3 ms | 260 ms cold | An agent that retrieves 50 times in a session waits ~0.6 s total, not 13 s — retrieval stops being the slow step in the loop. |
| — across all 41 datasets today | 701 ms | 15.3 ms | 260 ms | Our multi-index fan-out is serial today — measured, filed, fix scoped (#875: expected ~20–30 ms). |
| Exact phrase, p50 | 829 ms (fan-out) | 37.9 ms | 200 ms | Same #875 fan-out defect; the per-index engine is competitive once the query reaches it. |
| Regex over source | not offered | 870 ms | 180 ms | Nobody owns this today — Zoekt's regex loses to cold grep at this corpus size. A trigram side-car is on our roadmap precisely because the category is open. |
| Index 15 repos | 676 s + merge tail | 302 s | 0 s (no index) | XERJ's indexer does strictly more work — AST symbols in 13 languages and graph edges, not just trigrams — but the merge tail is a defect, not a feature (#876). |
| Idle daemon, this corpus | 8% CPU · 8.3 GB | 0.03% · 119 MB | none | Today's honest loss. Root-caused to three mechanisms with a published budget — idle under 0.5% of one core at any index count (#874). |
| Symbol / definition ranking | built in, 13 languages | needs universal-ctags installed | none | One binary gives your agent definition-first ranking with no sidecar toolchain to install or keep in sync. |
| Beyond code search | ES wire · aggregations · kNN vectors · agent memory | code search only | grep only | The same daemon that answers your agent's code lookups holds its logs, vectors and memory — one thing to run instead of three. |
Method: released XERJ binaries, Zoekt built from source at HEAD, ripgrep 14.1.1; identical
15-repo corpus (memcached, valkey, tantivy, regex, zstd, CRoaring, …), p50 of 21 requests per cell, CPU and RSS
from /proc. Zoekt indexed serially per its defaults; XERJ's indexer also extracts AST
symbols and graph edges, which Zoekt does not attempt — stated so the wall-clock rows are read fairly. Result
quality (ranking, symbol precision) is not scored here.
file:line
(#871–#876).
That is the deal this page offers: competitors are named, the red cells get the same precision as the green
ones, and each red cell links to its fix.The retrieval claim is not "search is nice" — it is that retrieval makes the same model
correct on APIs it has never memorised. Measured across 13 purpose-built libraries in 5 languages (unfamiliar
by construction, so the model cannot bluff), with hidden-test verdicts and real token accounting from
claude -p.
xerj autoindex parses source
through tree-sitter grammars in 13 languages and emits every symbol with its kind and line, a searchable
defs field, and the full body — so an identifier query ranks the definition
first. No ctags, no sidecar toolchain: one binary. Extraction throughput is measured and published
(~1,500 files/s single-thread on a 6,113-file Lucene checkout), and an unchanged re-index skips the parse
for byte-identical files (~100× on the edit-and-rerun path, shipped in rc.71).file:line: a few hundred tokens
carrying the one thing a compiler can never leak, the runtime contract.Method, per-library table, and every per-run record: the reference-coding case study · docs/case-studies/reference-coding
Two questions, both answered on public datasets anyone can download. How much does hybrid
retrieval add over BM25? And how many routing and spam decisions can labelled history settle with no model
call at all? Every vector and hybrid figure in this section was measured with
--embed-mode neural and the built-in all-MiniLM-L6-v2 model — not with the default
embedder, which is lexical feature hashing and has no model in it. The BM25 rows use no embedder and
hold on any node. The Jev rows below are written up verdict-first — wins, losses, and the judge's bill — in
the blog post: does XERJ beat JEV?
nDCG@10, BEIR test split | SciFact · 300 q · 5,183 docs | NFCorpus · 323 q · 3,633 docs | FiQA · 648 q · 57,638 docs | Run by |
|---|---|---|---|---|
XERJ BM25 (multi_match, no embedder) |
0.6572 | 0.3016 | 0.2382 | us · any node |
XERJ MiniLM vectors only (semantic) |
0.6764 | 0.3291 | not run | us · --embed-mode neural |
| XERJ BM25 top-30, reordered by MiniLM | 0.6855 | 0.3323 | not run | us · --embed-mode neural |
XERJ hybrid RRF (hybrid, server-side) |
0.6993 | 0.3448 | not run | us · --embed-mode neural |
XERJ BM25 top-30, reordered by Jev (jev-1.13.0, pinned) |
0.7410 | 0.3312 | 0.3638 | us · median of 3 repeats — the write-up |
| Jev (TypeSafe AI), hosted reranker | 0.768 | 0.358 | no figure | not run by us — published in the hev/jev-rerank README |
| Voyage rerank-3, hosted reranker | 0.755 | 0.357 | no figure | not run by us — same README |
| Cohere rerank-v3.5, hosted reranker | 0.745 | 0.340 | no figure | not run by us — same README |
rerank stage is now measured — and its first measurement was a
disaster we published. With an evaluation key granted by TypeSafe, the stage as it shipped scored
0.3822 nDCG@10 against BM25's 0.7750 on a 40-query pilot — our request asked the judge an
untargeted question, so it judged the pile once and echoed query-level noise. Root-caused, fixed
(0.8299, raw API 0.8389), then run in full on three datasets:
0.7410 / 0.3312 / 0.3638 over 1,271 judged queries, three repeats each. FiQA is the judge's
biggest lift (+0.126 over BM25) and its worst calibration case. The probabilities order documents well but
are not calibrated to BEIR relevance. ECE runs 0.10–0.31 across the three datasets, and
on FiQA documents rated 0.93 were relevant 34% of the time. Do not min_score
the raw number.
The full autopsy, run and gate. The stage is also the one XERJ feature
that sends document text off the machine —
read that before turning it on.| Decisions from labelled history · k = 10 · held-out items | Accuracy | Calibration error | Decided at confidence ≥ 0.8 | ms / item |
|---|---|---|---|---|
| Banking77, 77-way intent routing · 1,000 items — BM25, no model of any kind | 0.819 | 0.089 | 50.6% at 0.996 | 2.2 |
Banking77 — MiniLM neighbours (--embed-mode neural) |
0.933 | 0.012 | 86.9% at 0.979 | 8.6 |
Banking77 — hybrid RRF (--embed-mode neural) |
0.937 | 0.052 | 79.2% at 0.989 | 12.4 |
| SMS spam, yes / no · 1,000 items — BM25, no model of any kind | 0.983 | 0.017 | 95.4% at 0.992 | 0.9 |
SMS spam — MiniLM neighbours (--embed-mode neural) |
0.982 | 0.009 | 95.8% at 0.989 | 74.5 |
SMS spam — hybrid RRF (--embed-mode neural) |
0.989 | 0.015 | 95.7% at 0.997 | 67.2 |
| AG News, 4-way topic · 7,600 items · 120,000-deep history — BM25, no model of any kind | 0.9182 | 0.019 | 83.1% at 0.9669 | 34.2 |
POST /v1/systemone on the native port — so the pip-installed
jev-reranker runs against a local XERJ node unmodified, zero egress:
the write-up.Scripts, raw output and reproduce commands: benchmarks/beir-hybrid · benchmarks/decisions-as-retrieval · run 2026-09-18 on xerj v1.0.0-rc.74 · AG News 2026-09-20 on the wire-compat branch engine (cf6a9989) · written up in hybrid search quality and decisions without a model call. Hosted-reranker figures: hev/jev-rerank (MIT) · the measured Jev rows and the System One wire: blog/does-xerj-beat-jev.
Output tokens are the expensive kind — priced roughly 5× above input on Claude models — and
retry loops on unknown APIs burn exactly those. Three arms of the same agent on the same tasks: memory only,
grep-driven, retrieval-injected. Every figure from claude -p --output-format json.
The per-language medians run 6.7× (JavaScript) to 278× (Java) fewer output tokens than answering from memory. An independent 12-question dev-QA measurement on the engine's own reference corpora lands at 1.65× fewer output tokens and 1.47× cheaper — smaller, because it includes questions where the model already knew the answer; both numbers are published.
multi_match over defs/body/title, ~10 ms server-side.file:line — hundreds of tokens, not hundreds of thousands.The cheapest token is the one your model never generates. ZTA is the design rule that produced every number above: spend compute once, at index time, so agents stop spending inference tokens — the metered, per-request, 5×-priced resource — re-deriving what the index already knows. It is an architecture target with measured proxies, not a certification; here are its four principles and the number that keeps each one honest.
Parse the AST once — 13 languages, every symbol with kind and line, a ranked
defs field — instead of letting every future question re-derive structure with
tokens. Index-time compute is bought once; token-time compute is bought on every question, forever.
The unit of answer is the definition with its contract and file:line —
never "here are eleven files, good luck." Small answers keep the agent's context small, which compounds:
every later turn re-pays for everything already in the window.
Setup instructions live at llms.txt in machine order: install, start, index, query, retrieval discipline. One pasted sentence turns it on — tested verbatim, transcript published. Zero tokens spent negotiating with documentation written for humans.
An always-on corpus next to your agents must not tax the machine they work on. Zoekt sets the bar at 0.03% — the head-to-head above shows we are not there yet, and instead of hiding that, the remaining idle mechanisms are located, filed, and budgeted in public (#874). The rule the codebase now enforces: no per-index periodic work, ever.
Why this is the sales pitch and not a slogan: infrastructure used to compete on latency; agent infrastructure competes on your inference bill. A retrieval that answers in 10 ms and a few hundred tokens replaces a 15,000-token retry loop every time it fires — the $21.90 → $3.38 delta above is that substitution, measured 21 times over. ZTA is the commitment that every future XERJ feature is judged by the same question: how many tokens does it stop your model from spending?