THE CORPUS HUB: PINNED, LICENCED, SIGNED,
AND CONTRIBUTED BY PULL REQUEST
Every team that points an AI agent at a body of knowledge re-solves the same four problems in private: which exact version of the sources, under what licence, whether the copy you have is the copy I have, and who vouches for the bytes. We are opening the answers as a registry: the Corpus Hub, a branch on our GitHub where corpora arrive as pull requests — source pinned to SHAs, licences reviewed by a human per source, record packs checksummed and ed25519-signed. It ships with four reference corpora, one signed daily pack, a measured use case, and one result we did not want and are publishing anyway.
Update, 2026-10-06: the registry has outgrown this post. The
launch four are now 100 live corpora (119 backlog rows, every live one named in
the xerj-code skill's domain-selection table since rc.83) across seven categories —
including non-IT domains mined from real demand: eCFR titles 12/26/29 and UK legislation — and the
registry has its own site: browse every corpus with its licence verdict, pin, size, use case and
measured retrieval scores at hub.xerj.org.
What follows is the post as written at launch.
Agent context is everyone's private plumbing
Ask five teams what their coding agent retrieves from and you get five ad-hoc pipelines: a folder of cloned repos of unknown vintage, a wiki dump, a vector store someone seeded months ago. Each works, privately, and none of them can be shared, because a shared corpus needs four things a folder does not carry:
| Property | Why it matters | How the Hub does it |
|---|---|---|
| pinned | two people must retrieve the same bytes or results are incomparable | every repo recorded at the full SHA the build used; moving a pin is its own PR with a reason |
| licenced | "where did this text come from and may I act on it" is not optional the day someone pastes | a human review block per source — adapt-freely / read-don't-paste / mixed — not a detector's guess |
| checksummed | a corpus that silently drifts is a corpus that lies | packs carry per-file SHA-256; corpus add refuses the pack whole on mismatch |
| signed | checksums prove integrity; only a signature proves authorship | detached ed25519 over SHA256SUMS, verified with --verify-sig before a record is trusted |
That last row was the strangest finding of our pack work: of the vulnerability databases people pip-install and OCI-pull daily, none ships a signature a consumer can verify. The Hub does, because a corpus an agent reasons from is supply chain, and supply chain runs on authorship, not just integrity.
Two lanes: reference corpora and record packs
Reference corpora are other people's source code, cloned at pinned commits — the material for reference coding, where an agent reads how three other engines solved the problem before writing a fourth. Record packs are structured data you curate — the first is 1,950 Rust vulnerability advisories identity-resolved across osv.dev and RustSec, rebuilt and signed daily. One command each:
$ xerj corpus add --from https://raw.githubusercontent.com/xerj-org/xerj/corpus-hub/tools/xerj-code/hub/xerj-storage.json $ xerj corpus index xerj-storage $ xerj code xerj-storage "flush epoch crash recovery"
| corpus | lane | contents | use it for |
|---|---|---|---|
xerj-search | reference | lucene, tantivy, quickwit, meilisearch, sonic, elasticsearch | FTS, BM25, merge policy, ES wire semantics |
xerj-vector | reference | qdrant, usearch, instant-distance, hnswlib | HNSW construction, quantisation, filtered kNN |
xerj-storage | reference | sled, fjall, redb | WAL, crash recovery, compaction |
xerj-columnar | reference | clickhouse | columnar layout, codecs, vectorised scans |
rust-vulns | pack | 1,950 identity-resolved advisories, signed, daily | vulnerability precedent, affected-function lookup |
The grouping rule is the one design decision that matters: corpora are built by domain, not by language or popularity. A corpus that contains everything retrieves like a search engine with no query. The four reference corpora exist because an agent working on merge policy and an agent working on HNSW construction need different precedent — the same study that shows retrieval winning also shows the win gated by having the right precedent, not more of it.
Reference coding: the same solve rate, 2.7× fewer output tokens
The Hub's reason to exist is the measured one. In a controlled study — 16 tasks across five languages, on libraries written for the study so no training set contains them, each carrying a runtime contract the compiler cannot warn about — an agent with XERJ reference corpora and the same agent with native file access both solved 16/16. The difference was what it cost:
| native (grep + read files) | XERJ reference corpora | |
|---|---|---|
| tasks solved | 16/16 | 16/16 |
| output tokens (median) | 26,477 | 9,982 (2.7× fewer) |
| cost | $3.27 | $1.58 (2.1× cheaper) |
Read the fine print, because it is the same fine print the next section turns into a whole experiment: on the study's two memorised public controls (tantivy; valkey + memcached), retrieval was neutral-to-harmful — the model reproduces even a 256-value quantisation table from memory. The collapse happened exactly where the agent could not have known the answer — a runtime rule like "seal before you read", a recovery order no compiler will tell you. Without retrieval, the same model solved 1 of 21 such tasks and burned $21.9 flailing; with native grep but no retrieval focus, one corpus had ~1.06M input tokens of file content pulled into context.
The pack lane answers a differently-shaped question. Reviewing code and wondering whether a pattern has precedent:
$ xerj code rust-vulns "smallvec insert_many buffer overflow"
returns the identity-resolved advisory — one record carrying the RustSec prose, the GHSA filing, the CVE, the affected functions, the fixed versions — instead of three half-records in two databases that never talk to each other.
We A/B-tested the corpus on our own code. It tied.
Before opening a hub we ran the experiment that would justify it: two blinded arms auditing XERJ's own engine for security vulnerabilities — a plain agent, and the same agent with the rust-vulns corpus and XERJ retrieval. Same task, same turn budget, same pinned binary, two independent runs per arm, and a manual gate that re-proved every finding dynamically before counting it. Result, published in full:
| plain agent | corpus + XERJ | |
|---|---|---|
| findings filed | 20 | 18 |
| confirmed exploitable (dynamic PoC) | 13 | 16 |
| false positives | 1 — its only CRITICAL claim | 0 |
| cost / wall clock | $65.72 · 99.7 min | $70.55 · 89.8 min |
| findings traceable to the corpus | — | zero |
The corpus arm was marginally better on precision and it was never queried — one unused sweep call in one run, zero in the other. A tie, published as a tie, and then diagnosed against the pack itself: 1,963 records of advisory prose, only 14% containing any code, nothing that retrieves against a code-shaped question like "does slicing at a byte offset inside a multi-byte character have precedent?" — while half the pack was published after 2025, so memorisation was never the whole story. The failure was surface (nothing made retrieval the natural move), then shape (prose cannot match code), and only then size.
That diagnosis is now the Hub's roadmap, with pre-registered predictions: the next pack generation is code-shaped — version-pinned vulnerable functions extracted at the fix commit, before/after regression tests as trigger pairs, a bug-class taxonomy seeded from the 33 confirmed defect classes the two arms found — because we measured 9.8 extractable vulnerable functions per advisory and that is what a reviewer's question can actually retrieve against. If the re-test still ties, that gets published too. A registry whose numbers you can check is worth more than a demo.
One pull request, two lanes, a licence table
The Hub is a branch: corpus-hub
on xerj-org/xerj. Branch from it, open a PR against it, and a stdlib-only validator checks the
format on every push — manifest fields, full 40-hex SHAs, licence review blocks, recipe shape,
no binaries, no private keys. The full guide with templates lives in
/docs/corpus-hub and in the branch's
CONTRIBUTING.md.
What a reviewer will actually hold you to:
| rule | what it means |
|---|---|
| a nameable domain | finish the sentence "an agent working on ___ would query this for ___" or it is not ready |
| un-memorised content | retrieval's measured win is on niche, private, internal, post-cutoff material — not on famous snippets a model knows by heart |
| human licence review | a review block per source with one of three verdicts: adapt-with-attribution (Apache/MIT/BSD), approach-only (AGPL/SSPL/Elastic/BUSL/GPL — read the design, never paste), or mixed |
| no binaries in git | recipes and manifests only; built packs are signed Release assets a maintainer attaches after review |
The licence row is where we are least flexible, because we have been burned by our own tooling: our licence detector once reported Elasticsearch as Apache-2.0 — the AGPL text contains the phrase «an "Apache License 2.0" compatible license» — and a detector's false "permissive" is how restricted code ends up pasted into an Apache-2.0 project. So the CI checks that a human review block exists; only the reviewer's own opened licence file makes it true.
What is next for the Hub itself: the code-shaped pack generation above, and the pre-indexed half — a corpus you mount instead of indexing, so a consumer's CPU is not the price of entry. Both tracked in the open. If you have a domain worth retrieving from — engine internals, protocol specs, incident postmortems, hardware errata — the branch is open.