SIX SCARS FROM PRODUCTION
AGENT-MEMORY EVAL, FROM SOMEONE BLEEDING THERE.
We asked practitioners what a search engine built for agents is still missing. Chunxiao Wang of Assay (nautilus-compass) answered twice, in #1118 and #1138, with six patterns from a production agent-memory system. Each one cost a real failure, which is the only currency eval advice should trade in. This post is their material, credited, with two additions from our side: where their patterns rhyme with how this repo already gates itself, and the honest answer to the question they asked back, namely where our corpus registry is thin against their workload. It is thin in exactly the way they guessed, and we checked before saying so.
Six patterns, three themes: what changes when the corpus is the agent's own history, what an eval must freeze before anyone runs anything, and the arithmetic mistakes that make weak systems look strong. Every number in §01 through §04 is theirs. §05 is ours.
01 · THE CORPUS IS THE AGENT'S OWN HISTORY
Retrieving from a reference corpus and retrieving from session memory are different problems, and standard benchmarks cannot see the difference. Three of Chunxiao's observations carry it:
| The observation | Their measurement |
|---|---|
| the query is a situation, not a question | session-memory queries arrive as imperatives shaped by task state ("deploy failed again with 401"), not information needs ("how do 401s work?"). Hybrid dense+sparse retrieval handled these better than dense-only in their runs, but the bigger effect was chunking |
| topic slicing beat token-window slicing | chunking session memory by semantic topic instead of token window moved recall@1 by double digits in Chinese, and reversed the ranking of two embedding models relative to their MTEB scores. Benchmark-ledger performance did not predict in-distribution behavior on their session data |
| a failure hit is worth more than a manual hit | when the corpus is the agent's own memory, a hit that surfaces that you already tried this and it failed beats a hit that surfaces the manual page. Nobody scores that today, in their words |
And the failure mode that hides from every dashboard: a fill_diagonal-style masking bug that excludes gold pairs from the candidate pool turns a retrieval eval into a constant-zero or constant-pass exercise. They shipped one. The test costs nothing: if every query scores zero hits, suspect your mask, not your model.
02 · FOUR EVAL PATTERNS, EACH PAID FOR
The second contribution (#1118) is four patterns from building a verified-memory system. We keep them in their original order because the order is the story: distribution, leakage, calibration, variance.
| Pattern | Their measurement | The recipe |
|---|---|---|
| benchmark scores do not transfer to your traffic | bakeoff of bge-m3 vs Qwen3-embedding-0.6B on their own corpus (agent session memory, mixed zh/en): the MTEB-favored model gained +0.46pt recall@5 overall, under their transfer threshold, while the incumbent won recall@1 and the Chinese slices by +2.75pt | re-rank embedders on a frozen sample of your own traffic before switching; public benchmarks are priors, not verdicts |
| group by qid or your eval leaks | cross-item contamination (train items from the same task family appearing in eval retrieval context) silently inflated their early numbers | every eval split is grouped by question or task id, never by row |
| three-valued retrieval beats binary | their judge emits an unclear state with a confidence: Brier 0.069 with honest abstention, and the 2/27 abstentions in their latest public measurement were exactly the boundary rows |
let the retriever say don't know; score abstention as part of calibration (ECE), not as failure |
| to measure system drift, cache the inputs | same-day score swings of 0.93 → 0.30 → 0.90 looked like model drift; replaying cached inputs through the judge K times showed 0.0pp judge-side variance, all of it in the generated side | two-stage variance attribution, replay-cached-input vs regenerate-then-judge, before blaming retrieval or judging |
03 · PREREGISTER, LABEL, BASELINE, PAIR
The third theme (#1138, recipe 2) is the discipline around the run itself, stated in four rules:
| The rule | The scar behind it |
|---|---|
| freeze criteria before running: metric, threshold, stop-loss, and the exact verifiable artifact each claim points to. Criteria may only move stricter mid-stream; loosening needs a fresh preregistration | every rule here is a re-learned lesson from a run that was allowed to move its own goalposts |
| label every claim measured, inferred, or unverifiable. Inferred claims carry an upgrade path | the most useful eval artifact they ever published was a self-report reading "2 of 4 gates FAIL on recompute". The visible FAILs are what made the green credible afterwards |
| compute the majority-class baseline before claiming accuracy | a 0.87 binary classifier that a 0.78 always-yes baseline nearly matched, n=60, confidence intervals overlapping. Below n=100 the CI belongs next to every headline number, or the number is not a claim, it is a mood |
| paired designs beat unpaired by an order of magnitude of information | same frames, two checkpoints, McNemar. Their unpaired reading overstated an effect by roughly 4× versus the paired one |
The affinity is not an accident. This repo gates its own releases the same way: bars written before the work, FAILs kept visible in the notes, and a CI job that re-checks every claim in a release section against the tree it describes. The rc.80 gates post is the same philosophy with our numbers in it: three gates passed, three failed, one split, all published. When Chunxiao writes that visible FAILs are what made the green credible, that is the operating manual we already live by, arriving independently from a different domain. That convergence is worth more than either set of numbers alone.
04 · THE QUESTION THEY ASKED BACK, ANSWERED HONESTLY
Their price for the recipes was not a star. It was this: an honest issue pointing at where the index is thin against these patterns, especially session-memory slicing and failure-turn retention. We checked the registry rather than answering from memory. The answer, posted in full in #1138:
Three thin spots, stated plainly. First, the corpus they described, real multi-session agent trajectories with timestamps and failure turns kept verbatim, does not exist at our hub: zero of the 99 live manifests and none of the other 131 backlog entries are session or trajectory data. Nobody using our reference corpora can score what they score. Second, our indexer chunks at file and section boundaries and keeps a locator back to the source span; for session logs that is the wrong shape, and we have no semantic-topic slicer to offer instead. Third, our memory features store curated facts an agent chose to keep, with BM25 recall by default; there is no native tried-and-failed marker and no retention of turns that went nowhere.
What we did about it: the wanted corpus now has a named slot in the hub backlog (agent-session-trajectories, status planned, demand-anchored to their issue), and if Assay can donate sanitized multi-session trajectories with failure turns intact, the contribution path is open by pull request. It would be a first-of-kind corpus at the hub. Their offer to recompute our benchmarks independently stands unaccepted for now, and it should not: an independent recompute of any number we publish is always welcome.
If you have production scars of your own: the same invitation stands. Open an issue with what it cost you and what you measured, and if it survives our check against reality, it gets the same treatment, your name on it. The engine itself is one binary with a quickstart, and the agent-memory recipe is the closest thing we ship today to what §01 describes wanting.
Guest material verbatim in issues #1118 and #1138 · their repository: nautilus-compass · the honest-gap comment: #1138 comment · backlog entry PR #1210 · related: the rc.80 gates · the Corpus Hub · the agent-memory recipe.