CORPUS HUB 01Registry & contributing
01 · CORPUS HUB

Corpus Hub: registry & contributing

The Corpus Hub is a public registry of corpora for AI agents: reference source code and structured record packs, each pinned to exact commits, licence-reviewed by a human, and — for packs — checksummed and signed. It lives on the corpus-hub branch of xerj-org/xerj, and contributions arrive as pull requests against that branch. Browse the whole registry as a site — every corpus with its category, use case, licence verdict, pin, size and measured retrieval scores — at hub.xerj.org; that site is generated from the branch, so it is never stale relative to the registry. This page is the walkthrough; the branch's own CONTRIBUTING.md is the source of record for the rules.

Source · tools/xerj-code/hub/CONTRIBUTING.md (corpus-hub branch) · launch post: /blog/the-corpus-hub · browse: hub.xerj.org

Consuming a corpus

Three commands, no build step — the binary indexes from the pinned manifest directly:

$ xerj corpus add --from https://raw.githubusercontent.com/xerj-org/xerj/corpus-hub/tools/xerj-code/hub/xerj-storage.json
$ xerj corpus index xerj-storage
$ xerj code xerj-storage "flush epoch crash recovery"

Clones live under ~/.xerj-code/corpora/, the index under ~/.xerj-code/data — both persist across reboots. An index older than 30 days is refused by design: a stale index is worse than none, so re-run xerj corpus index rather than forcing a stale answer. Pack consumers additionally verify the detached ed25519 signature over SHA256SUMS before trusting a record.

Two lanes

Every entry in the Hub is one of two kinds, and the lane decides the format you submit:

LANE A — reference corpus   tools/xerj-code/hub/<name>.json   (a manifest)
LANE B — record pack        tools/packs/<name>/recipe.toml     (a recipe)

A reference corpus is other people's source code at pinned commits — the material for reference coding ("how do three other engines implement merge policy?"). A record pack is structured data you curate and resolve — the first is 1,950 Rust vulnerability advisories identity-merged across osv.dev and RustSec, rebuilt and signed daily. If your content is "files someone can clone", it is Lane A; if it is "records distilled from one or more sources", it is Lane B.

Lane A — the manifest

Copy tools/xerj-code/hub/TEMPLATE.json to tools/xerj-code/hub/<your-corpus>.json. The corpus name must equal the filename. Field by field:

{
  "corpus": "xerj-storage",
  "cloned_at": "<UTC timestamp the build ran>",
  "repos": [
    {
      "repo": "sled",
      "url": "https://github.com/spacejam/sled",
      "licence": "<what the detector read — do not trust this alone>",
      "sha": "<full 40-hex commit the build actually used>",
      "files": 417, "bytes": 2900000,
      "review": {
        "spdx": "<SPDX id YOU confirmed by opening the licence file>",
        "use": "adapt-with-attribution",
        "by": "your-handle", "at": "2026-10-01",
        "note": "<one line: what you opened, what you concluded>"
      }
    }
  ]
}

The two fields CI cannot check, and a reviewer will hold you to: sha must be the commit the build actually used (never a branch head), and review must be a human's verdict — see the licence contract below.

Lane B — the recipe

Copy tools/packs/TEMPLATE-recipe.toml to tools/packs/<name>/recipe.toml. Everything domain-specific lives in the recipe, never in the tooling:

[recipe]        format = 1 · name = directory name · description
[[sources]]     one per source: slug, kind (git|http-zip|dir), url,
                glob, format (flat|osv|rustsec-md), licence
[envelope]      how a raw source file becomes one searchable record:
                id_from, title_from, body_join, defs_from
[identity]      which fields make two records the same thing
                (edges, each=true explodes list fields; union-find
                closes the chain) + canonical_source_order
[merge]         precedence when two sources carry the same fact
[[derived]]     optional build-time fields (regex_extract, present)
[emit]          shards

The identity block is the whole dedup contract — write it so a reader can predict the merged count, then put the measured count (envelopes → records) in your README. The showcase to read alongside the template is tools/packs/rust-vulns/recipe.toml: every block there carries a comment saying why the source is in and what it was measured to uniquely contribute.

Built packs are never committed. A maintainer builds from your recipe, attaches the signed pack to a GitHub Release, and the manifest links it — the registry in git stays recipes-and-manifests only.

The licence contract

Every source carries a human review block with one of three verdicts:

adapt-with-attribution   Apache-2.0 · MIT · BSD — cite file:line when
                         you adapt; copying is fine within the licence
approach-only            AGPL · SSPL · Elastic · BUSL · GPL — read the
                         design, write your own code, never paste
mixed                    per-file split (e.g. MIT core + BUSL edition) —
                         name the boundary in the note

This is the rule we are least flexible on, because we have been burned by our own tooling: our licence detector once reported Elasticsearch as Apache-2.0 — the AGPL text contains the phrase «an "Apache License 2.0" compatible license». A detector's false "permissive" is exactly how restricted code ends up pasted into an Apache-2.0 project. CI checks that a review block exists and its fields are well-formed; only the reviewer's own opened licence file makes it true.

The pull request

# 1. branch FROM corpus-hub (not main) so the CI workflow is in your PR
  git clone -b corpus-hub https://github.com/xerj-org/xerj && cd xerj

# 2. add your manifest or recipe, then validate locally (stdlib only)
  python3 tools/xerj-code/hub/validate_hub.py

# 3. commit, push, open the PR against corpus-hub
  git checkout -b corpus/<your-name>
  git add tools/xerj-code/hub/<your-name>.json
  git commit -m "corpus: add <your-name> (<domain>)"
  git push -u origin corpus/<your-name>

CI on every push and PR runs the same validator plus two registry guards: no file over 5 MB under the registry paths (built packs and key material must not be in git), and no private key material under tools/packs/keys. A maintainer then reviews — domain fit, pin truthfulness, and above all the licence verdicts — and merges. After merge, the 30-day freshness refusal applies to your corpus like every other: keeping it current is part of contributing it.

Is your corpus a fit?

Two questions, both measured, not vibes:

1. Can you name the domain? Finish the sentence: "an agent working on ___ would query this for ___." Corpora are built by domain, not by language or popularity — a corpus that contains everything retrieves like a search engine with no query.

2. Is the content un-memorised? Retrieval's measured win is on niche, private, internal, post-cutoff material. On famous public library code a model already knows, retrieval is overhead — our own study showed it saving nothing on the two memorised tasks. Do not force it where it loses.