← ALL POSTS · /BLOG

CYBER-EXPLOITS CORPUS · A/B EVAL · SEALED GROUND TRUTH · RUN 2026-10-08

THE EXPLOIT HUB, TESTED ON
THREE FRESH CVES

Thirty-six proof-of-concept repositories for CVE-2026-23744 sit in our exploit corpus. One query over 1.09 million PoC documents named the exact route, the operative field, and the one-request trigger, with citations. The advisory alone could not name any of the three. Then we unsealed the vendor patch, and both arms, with the corpus and without, had predicted the wrong fix. That is the whole eval in one sentence: the corpus answered attribution questions nothing else could, and did not magically answer the one nobody could. This post is the full run, losses included.

01 · THE DESIGN

Two arms, sealed ground truth, and a negative control we ran against ourselves

We picked three CVEs published after the corpus pin date and ran each through two arms of the same analyst session:

Ground truth was sealed before Arm B wrote a word: vendor patches downloaded and diffed unread, PoC repositories cloned but unopened, then unsealed only after both arms were frozen. Every query Arm B ran is logged in the artifacts, including the ones that failed.

Three honesty rules shaped the scoring. First, this is a pilot: n=3, one analyst, so it is a case study with measured arms, not statistics. Second, when retrieval adds nothing we count the case as added-nothing rather than dropping it from the denominator. Third, we ran a negative control against our own claim: cloning the three needle PoC repos and grepping them took 2.1 seconds and found the same endpoint, with the repo names already in hand. The corpus is not a fetch accelerator. It is the discovery layer: Arm A had no way to know those repositories existed, which three to read, or that 33 others corroborate the same route.

02 · THE NEEDLE

CVE-2026-23744: one query, exact attribution, wrong patch

The one case whose PoCs are in the corpus. The question a working analyst asks first is not “what class is this?” but “what exactly do I send, and where?”

QuestionArm A (advisory)Arm B (corpus)Sealed truth
technique classright, inferred from CWE-306 proseright, cited from 3 independent PoCsunauth endpoint + stdio command launch
exact route + fieldcould not namePOST /api/mcp/connect, serverConfig.commandconfirmed in all PoCs
trigger stepsambiguous (install vs connect)one POST, command launches in-requestconfirmed
patch mechanism3 candidates, argued auth+transportsame pair, argued more confidentlybind 127.0.0.1 only; both arms wrong
patch filecould not nameguessed the route fileserver/index.ts (hostname constant)

The retrieval cost was measured, not vibes: two shipped-client queries (14.2 s and 5.7 s, both failed for reasons in section 05) plus three direct multi_match queries at about a second each, roughly 23 seconds end to end over 1.09M documents. The vendor patch, when unsealed, turned out to be one line: the stdio server hostname constant moved from 0.0.0.0 to 127.0.0.1. The unauthenticated route is still there. Both arms dismissed that candidate as insufficient, and the corpus evidence made Arm B more confident in the wrong pair. Attribution and fix prediction are different questions, and the corpus answered only the one it had evidence for.

03 · THE ANALOGY

CVE-2026-105844: no PoC anywhere, so the corpus lent a neighbor

No PoC for this CVE exists in the corpus or, at eval time, anywhere public. Arm B ran on analogy retrieval instead: how do same-year prototype-pollution exploits in the corpus actually work? The productive query returned an in-corpus 2026 writeup (exploitdb entry 52528, deephas 1.0.7, CVE-2026-25047) that names both path shapes and the filter-bypass variants:

"1. constructor.prototype path + hasOwnProperty bypass
 2. __proto__ path + indexOf bypass"
gadget family: process.env, require.extensions, child_process
       or: hasOwnProperty / toString pollution for auth bypass

That transformed the patch prediction from a shape into a spec. Arm B predicted a denylist of prototype-sensitive path segments plus a null-prototype accumulator. The sealed 3.88.0 diff shipped a new fieldPath.ts with Set(['__proto__','constructor','prototype']) and an Object.create(null) in the nested-assignment helper, plus a 400 on bad paths. Both mechanisms named, the denylist contents named. The losses: neither arm could name the vulnerable file (it was setNestedValue.ts and friends in the import-export plugin), and Arm B narrowed the entry point to the import half when the truth was the export preview. Arm A had hedged both halves and was right by hedging. Analogies sharpen mechanisms and can over-narrow entry points.

04 · THE EMPTY STRATUM

CVE-2026-97332: the case the corpus added nothing, counted

A WordPress multisite plugin whose private uploads directory is served directly because the root-level rewrite rule never matches the per-site path. The advisory title already states the mechanism, so Arm A was concrete from prose alone. Arm B ran three queries against the exploitdb stratum looking for a close precedent for the class and found none: that stratum is weighted to older RCE/SQLi/upload classes, and WPScan-sourced CVEs rarely get GitHub PoCs at all.

Retrieval cannot manufacture what is not indexed. The case is counted as adds ~0 in the final tally rather than dropped, because a corpus eval that only reports its hits is marketing. The sealed patch was a fourth shape nobody listed: a folder-level .htaccess inside the uploads directory, generated at activation, with a Require all denied fallback for hosts without mod_rewrite.

The tally across the pilot: technique attribution 2 of 3 improved by the corpus (one from nothing to exact, one from generic to same-year-cited), patch-mechanism prediction 1 of 3 improved, 1 of 3 cases added nothing. Two related measurements fell out of the run: zero of 128 post-pin CVEs checked had any public PoC at eval time, so fresh-CVE PoC evaluation needs either pin-rolling or corpus-present CVEs as the eval set; and the raw group stands at about 37.6 GB against the 50 GB build target, a number we report as it is because the gap is fetch backlog, not padding.

05 · THE DEFECT

The eval caught a live bug in our own client

Arm B's first two queries used the shipped xerj code client and surfaced zero of the 36 indexed needle PoCs on natural queries. That is not acceptable in a product whose pitch is exactly this query, so we dug in instead of routing around it.

Mechanism, measured: the client builds its field list from the union of mappings across every index in the corpus. One index family in the group maps an extra text field, so the union grows to four fields, and the engine then zeroes every multi-token multi_match for the roughly five thousand indices that do not map the fourth field. Drop the poison field and the needle appears:

multi_match field set (same query, same corpus)HitsNeedle PoCs in top results
[body, defs, title, text] (client union, shipped)39absent; one 23,745-word wordlist index dominates
[body, defs, title]2716 distinct needle docs in the top 27

Filed as issue #1238 with the table above, the fix direction (unmapped fields should be skipped per index with a hint, the way search engines with mappings already treat them, and the client should intersect mappings rather than union them), and a two-index regression test proposal. We are stating the obvious conclusion: an A/B harness run over a real corpus doubles as corpus QA, and the eval paid for itself before the blog post existed.

06 · WHAT IS IN THE GROUP

Four strata, one join, and a licence line we do not blur

StratumContentsStatus at eval time
exploit-pocs-2026PoC repository source, per-repo indices, full-text1,093,749 docs / 5,644 indices, live
exploitdb-srcExploit-DB entries with source, full-text218,329 docs, live
cve-recordsCVE/NVD records with CVSS, references, CNA prosepartial, apply phase in progress
vuln-fix-commitsthe fix side: inventoried commits linking advisories to patches8,546 / 30,207 payloads fetched, backfill running

The join the group exists for is the one nobody ships: advisory, working exploit source, and the fix commit in one searchable place, so an agent can ask “what does exploitation of this class actually look like this year” and cite the answer. Licence posture is deliberate and one line: Exploit-DB is GPL and PoC repositories carry their own terms, so the whole group is approach-only, indexed for retrieval, never copied into XERJ and never redistributed as a pack. Query it, cite it, write your own code.

07 · THE ADD-ON

A corpus of our own posts, because agents write too

One more corpus shipped alongside this eval, and it is the smallest and strangest one on the hub: xerj-blogposts, seven posts (this one included at the next rebuild) as retrievable records with derived CVE and release ids, 107,295 bytes, built from the published pages by a generator whose --check mode fails CI if the corpus drifts from the site. It exists because the writing task is also a retrieval task: an agent asked to write the next XERJ post needs the house voice as data, how these posts open (a number first), argue (every claim traced to a run), and close (the losses left in). This post is the proof by construction; it was written the way it argues, with the sealed patches, the failed queries, the 2.1-second negative control, and the empty-stratum case all in the tables, and the corpus that taught it the shape is now on the hub beside the exploit group it reports on.