Corpus Hub: registry & contributing
The Corpus Hub is a public registry of corpora for AI agents: reference source code and
structured record packs, each pinned to exact commits, licence-reviewed by a human, and — for
packs — checksummed and signed. It lives on the
corpus-hub
branch of xerj-org/xerj, and contributions arrive as pull requests against that branch.
Browse the whole registry as a site — every corpus with its category, use case, licence verdict,
pin, size and measured retrieval scores — at hub.xerj.org;
that site is generated from the branch, so it is never stale relative to the registry.
This page is the walkthrough; the branch's own
CONTRIBUTING.md
is the source of record for the rules.
Source · tools/xerj-code/hub/CONTRIBUTING.md (corpus-hub branch) · launch post: /blog/the-corpus-hub · browse: hub.xerj.org
Consuming a corpus
Three commands, no build step — the binary indexes from the pinned manifest directly:
$ xerj corpus add --from https://raw.githubusercontent.com/xerj-org/xerj/corpus-hub/tools/xerj-code/hub/xerj-storage.json $ xerj corpus index xerj-storage $ xerj code xerj-storage "flush epoch crash recovery"
Clones live under ~/.xerj-code/corpora/, the index under
~/.xerj-code/data — both persist across reboots. An index older than
30 days is refused by design: a stale index is worse than none, so re-run
xerj corpus index rather than forcing a stale answer. Pack consumers
additionally verify the detached ed25519 signature over SHA256SUMS before trusting a record.
Two lanes
Every entry in the Hub is one of two kinds, and the lane decides the format you submit:
LANE A — reference corpus tools/xerj-code/hub/<name>.json (a manifest) LANE B — record pack tools/packs/<name>/recipe.toml (a recipe)
A reference corpus is other people's source code at pinned commits — the material for reference coding ("how do three other engines implement merge policy?"). A record pack is structured data you curate and resolve — the first is 1,950 Rust vulnerability advisories identity-merged across osv.dev and RustSec, rebuilt and signed daily. If your content is "files someone can clone", it is Lane A; if it is "records distilled from one or more sources", it is Lane B.
Lane A — the manifest
Copy tools/xerj-code/hub/TEMPLATE.json to
tools/xerj-code/hub/<your-corpus>.json. The corpus name must equal
the filename. Field by field:
{
"corpus": "xerj-storage",
"cloned_at": "<UTC timestamp the build ran>",
"repos": [
{
"repo": "sled",
"url": "https://github.com/spacejam/sled",
"licence": "<what the detector read — do not trust this alone>",
"sha": "<full 40-hex commit the build actually used>",
"files": 417, "bytes": 2900000,
"review": {
"spdx": "<SPDX id YOU confirmed by opening the licence file>",
"use": "adapt-with-attribution",
"by": "your-handle", "at": "2026-10-01",
"note": "<one line: what you opened, what you concluded>"
}
}
]
}
The two fields CI cannot check, and a reviewer will hold you to: sha
must be the commit the build actually used (never a branch head), and
review must be a human's verdict — see
the licence contract below.
Lane B — the recipe
Copy tools/packs/TEMPLATE-recipe.toml to
tools/packs/<name>/recipe.toml. Everything domain-specific lives
in the recipe, never in the tooling:
[recipe] format = 1 · name = directory name · description
[[sources]] one per source: slug, kind (git|http-zip|dir), url,
glob, format (flat|osv|rustsec-md), licence
[envelope] how a raw source file becomes one searchable record:
id_from, title_from, body_join, defs_from
[identity] which fields make two records the same thing
(edges, each=true explodes list fields; union-find
closes the chain) + canonical_source_order
[merge] precedence when two sources carry the same fact
[[derived]] optional build-time fields (regex_extract, present)
[emit] shards
The identity block is the whole dedup contract — write it so a reader can predict the merged
count, then put the measured count (envelopes → records) in your README. The showcase to read
alongside the template is tools/packs/rust-vulns/recipe.toml: every
block there carries a comment saying why the source is in and what it was measured to uniquely
contribute.
Built packs are never committed. A maintainer builds from your recipe, attaches the signed pack to a GitHub Release, and the manifest links it — the registry in git stays recipes-and-manifests only.
The licence contract
Every source carries a human review block with one of three verdicts:
adapt-with-attribution Apache-2.0 · MIT · BSD — cite file:line when
you adapt; copying is fine within the licence
approach-only AGPL · SSPL · Elastic · BUSL · GPL — read the
design, write your own code, never paste
mixed per-file split (e.g. MIT core + BUSL edition) —
name the boundary in the note
This is the rule we are least flexible on, because we have been burned by our own tooling: our licence detector once reported Elasticsearch as Apache-2.0 — the AGPL text contains the phrase «an "Apache License 2.0" compatible license». A detector's false "permissive" is exactly how restricted code ends up pasted into an Apache-2.0 project. CI checks that a review block exists and its fields are well-formed; only the reviewer's own opened licence file makes it true.
The pull request
# 1. branch FROM corpus-hub (not main) so the CI workflow is in your PR git clone -b corpus-hub https://github.com/xerj-org/xerj && cd xerj # 2. add your manifest or recipe, then validate locally (stdlib only) python3 tools/xerj-code/hub/validate_hub.py # 3. commit, push, open the PR against corpus-hub git checkout -b corpus/<your-name> git add tools/xerj-code/hub/<your-name>.json git commit -m "corpus: add <your-name> (<domain>)" git push -u origin corpus/<your-name>
CI on every push and PR runs the same validator plus two registry guards: no file over 5 MB
under the registry paths (built packs and key material must not be in git), and no private key
material under tools/packs/keys. A maintainer then reviews —
domain fit, pin truthfulness, and above all the licence verdicts — and merges. After merge, the
30-day freshness refusal applies to your corpus like every other: keeping it current is part of
contributing it.
Is your corpus a fit?
Two questions, both measured, not vibes:
1. Can you name the domain? Finish the sentence: "an agent working on ___ would query this for ___." Corpora are built by domain, not by language or popularity — a corpus that contains everything retrieves like a search engine with no query.
2. Is the content un-memorised? Retrieval's measured win is on niche, private, internal, post-cutoff material. On famous public library code a model already knows, retrieval is overhead — our own study showed it saving nothing on the two memorised tasks. Do not force it where it loses.