← ALL POSTS · /BLOG

CORPUS HUB · PINNED · LICENCED · SIGNED · BUILT 2026-10-01

THE CORPUS HUB: PINNED, LICENCED, SIGNED,
AND CONTRIBUTED BY PULL REQUEST

Every team that points an AI agent at a body of knowledge re-solves the same four problems in private: which exact version of the sources, under what licence, whether the copy you have is the copy I have, and who vouches for the bytes. We are opening the answers as a registry: the Corpus Hub, a branch on our GitHub where corpora arrive as pull requests — source pinned to SHAs, licences reviewed by a human per source, record packs checksummed and ed25519-signed. It ships with four reference corpora, one signed daily pack, a measured use case, and one result we did not want and are publishing anyway.

Update, 2026-10-06: the registry has outgrown this post. The launch four are now 100 live corpora (119 backlog rows, every live one named in the xerj-code skill's domain-selection table since rc.83) across seven categories — including non-IT domains mined from real demand: eCFR titles 12/26/29 and UK legislation — and the registry has its own site: browse every corpus with its licence verdict, pin, size, use case and measured retrieval scores at hub.xerj.org. What follows is the post as written at launch.

01 · THE PROBLEM

Agent context is everyone's private plumbing

Ask five teams what their coding agent retrieves from and you get five ad-hoc pipelines: a folder of cloned repos of unknown vintage, a wiki dump, a vector store someone seeded months ago. Each works, privately, and none of them can be shared, because a shared corpus needs four things a folder does not carry:

PropertyWhy it mattersHow the Hub does it
pinnedtwo people must retrieve the same bytes or results are incomparableevery repo recorded at the full SHA the build used; moving a pin is its own PR with a reason
licenced"where did this text come from and may I act on it" is not optional the day someone pastesa human review block per source — adapt-freely / read-don't-paste / mixed — not a detector's guess
checksummeda corpus that silently drifts is a corpus that liespacks carry per-file SHA-256; corpus add refuses the pack whole on mismatch
signedchecksums prove integrity; only a signature proves authorshipdetached ed25519 over SHA256SUMS, verified with --verify-sig before a record is trusted

That last row was the strangest finding of our pack work: of the vulnerability databases people pip-install and OCI-pull daily, none ships a signature a consumer can verify. The Hub does, because a corpus an agent reasons from is supply chain, and supply chain runs on authorship, not just integrity.

02 · WHAT IS IN IT

Two lanes: reference corpora and record packs

Reference corpora are other people's source code, cloned at pinned commits — the material for reference coding, where an agent reads how three other engines solved the problem before writing a fourth. Record packs are structured data you curate — the first is 1,950 Rust vulnerability advisories identity-resolved across osv.dev and RustSec, rebuilt and signed daily. One command each:

$ xerj corpus add --from https://raw.githubusercontent.com/xerj-org/xerj/corpus-hub/tools/xerj-code/hub/xerj-storage.json
$ xerj corpus index xerj-storage
$ xerj code xerj-storage "flush epoch crash recovery"
corpuslanecontentsuse it for
xerj-searchreferencelucene, tantivy, quickwit, meilisearch, sonic, elasticsearchFTS, BM25, merge policy, ES wire semantics
xerj-vectorreferenceqdrant, usearch, instant-distance, hnswlibHNSW construction, quantisation, filtered kNN
xerj-storagereferencesled, fjall, redbWAL, crash recovery, compaction
xerj-columnarreferenceclickhousecolumnar layout, codecs, vectorised scans
rust-vulnspack1,950 identity-resolved advisories, signed, dailyvulnerability precedent, affected-function lookup

The grouping rule is the one design decision that matters: corpora are built by domain, not by language or popularity. A corpus that contains everything retrieves like a search engine with no query. The four reference corpora exist because an agent working on merge policy and an agent working on HNSW construction need different precedent — the same study that shows retrieval winning also shows the win gated by having the right precedent, not more of it.

03 · THE USE CASE, MEASURED

Reference coding: the same solve rate, 2.7× fewer output tokens

The Hub's reason to exist is the measured one. In a controlled study — 16 tasks across five languages, on libraries written for the study so no training set contains them, each carrying a runtime contract the compiler cannot warn about — an agent with XERJ reference corpora and the same agent with native file access both solved 16/16. The difference was what it cost:

native (grep + read files)XERJ reference corpora
tasks solved16/1616/16
output tokens (median)26,4779,982 (2.7× fewer)
cost$3.27$1.58 (2.1× cheaper)

Read the fine print, because it is the same fine print the next section turns into a whole experiment: on the study's two memorised public controls (tantivy; valkey + memcached), retrieval was neutral-to-harmful — the model reproduces even a 256-value quantisation table from memory. The collapse happened exactly where the agent could not have known the answer — a runtime rule like "seal before you read", a recovery order no compiler will tell you. Without retrieval, the same model solved 1 of 21 such tasks and burned $21.9 flailing; with native grep but no retrieval focus, one corpus had ~1.06M input tokens of file content pulled into context.

The pack lane answers a differently-shaped question. Reviewing code and wondering whether a pattern has precedent:

$ xerj code rust-vulns "smallvec insert_many buffer overflow"

returns the identity-resolved advisory — one record carrying the RustSec prose, the GHSA filing, the CVE, the affected functions, the fixed versions — instead of three half-records in two databases that never talk to each other.

04 · THE MEASUREMENT WE DID NOT WANT

We A/B-tested the corpus on our own code. It tied.

Before opening a hub we ran the experiment that would justify it: two blinded arms auditing XERJ's own engine for security vulnerabilities — a plain agent, and the same agent with the rust-vulns corpus and XERJ retrieval. Same task, same turn budget, same pinned binary, two independent runs per arm, and a manual gate that re-proved every finding dynamically before counting it. Result, published in full:

plain agentcorpus + XERJ
findings filed2018
confirmed exploitable (dynamic PoC)1316
false positives1 — its only CRITICAL claim0
cost / wall clock$65.72 · 99.7 min$70.55 · 89.8 min
findings traceable to the corpus—zero

The corpus arm was marginally better on precision and it was never queried — one unused sweep call in one run, zero in the other. A tie, published as a tie, and then diagnosed against the pack itself: 1,963 records of advisory prose, only 14% containing any code, nothing that retrieves against a code-shaped question like "does slicing at a byte offset inside a multi-byte character have precedent?" — while half the pack was published after 2025, so memorisation was never the whole story. The failure was surface (nothing made retrieval the natural move), then shape (prose cannot match code), and only then size.

That diagnosis is now the Hub's roadmap, with pre-registered predictions: the next pack generation is code-shaped — version-pinned vulnerable functions extracted at the fix commit, before/after regression tests as trigger pairs, a bug-class taxonomy seeded from the 33 confirmed defect classes the two arms found — because we measured 9.8 extractable vulnerable functions per advisory and that is what a reviewer's question can actually retrieve against. If the re-test still ties, that gets published too. A registry whose numbers you can check is worth more than a demo.

05 · CONTRIBUTE

One pull request, two lanes, a licence table

The Hub is a branch: corpus-hub on xerj-org/xerj. Branch from it, open a PR against it, and a stdlib-only validator checks the format on every push — manifest fields, full 40-hex SHAs, licence review blocks, recipe shape, no binaries, no private keys. The full guide with templates lives in /docs/corpus-hub and in the branch's CONTRIBUTING.md. What a reviewer will actually hold you to:

rulewhat it means
a nameable domainfinish the sentence "an agent working on ___ would query this for ___" or it is not ready
un-memorised contentretrieval's measured win is on niche, private, internal, post-cutoff material — not on famous snippets a model knows by heart
human licence reviewa review block per source with one of three verdicts: adapt-with-attribution (Apache/MIT/BSD), approach-only (AGPL/SSPL/Elastic/BUSL/GPL — read the design, never paste), or mixed
no binaries in gitrecipes and manifests only; built packs are signed Release assets a maintainer attaches after review

The licence row is where we are least flexible, because we have been burned by our own tooling: our licence detector once reported Elasticsearch as Apache-2.0 — the AGPL text contains the phrase «an "Apache License 2.0" compatible license» — and a detector's false "permissive" is how restricted code ends up pasted into an Apache-2.0 project. So the CI checks that a human review block exists; only the reviewer's own opened licence file makes it true.

What is next for the Hub itself: the code-shaped pack generation above, and the pre-indexed half — a corpus you mount instead of indexing, so a consumer's CPU is not the price of entry. Both tracked in the open. If you have a domain worth retrieving from — engine internals, protocol specs, incident postmortems, hardware errata — the branch is open.