← ALL POSTS · /BLOG

GUEST CONTRIBUTION · AGENT MEMORY · RETRIEVAL EVAL · ISSUES #1118 + #1138

SIX SCARS FROM PRODUCTION
AGENT-MEMORY EVAL, FROM SOMEONE BLEEDING THERE.

WHY THIS POST EXISTS a guest answered our invitation with receipts

We asked practitioners what a search engine built for agents is still missing. Chunxiao Wang of Assay (nautilus-compass) answered twice, in #1118 and #1138, with six patterns from a production agent-memory system. Each one cost a real failure, which is the only currency eval advice should trade in. This post is their material, credited, with two additions from our side: where their patterns rhyme with how this repo already gates itself, and the honest answer to the question they asked back, namely where our corpus registry is thin against their workload. It is thin in exactly the way they guessed, and we checked before saying so.

Six patterns, three themes: what changes when the corpus is the agent's own history, what an eval must freeze before anyone runs anything, and the arithmetic mistakes that make weak systems look strong. Every number in §01 through §04 is theirs. §05 is ours.

01 · THE CORPUS IS THE AGENT'S OWN HISTORY

Retrieving from a reference corpus and retrieving from session memory are different problems, and standard benchmarks cannot see the difference. Three of Chunxiao's observations carry it:

The observationTheir measurement
the query is a situation, not a question session-memory queries arrive as imperatives shaped by task state ("deploy failed again with 401"), not information needs ("how do 401s work?"). Hybrid dense+sparse retrieval handled these better than dense-only in their runs, but the bigger effect was chunking
topic slicing beat token-window slicing chunking session memory by semantic topic instead of token window moved recall@1 by double digits in Chinese, and reversed the ranking of two embedding models relative to their MTEB scores. Benchmark-ledger performance did not predict in-distribution behavior on their session data
a failure hit is worth more than a manual hit when the corpus is the agent's own memory, a hit that surfaces that you already tried this and it failed beats a hit that surfaces the manual page. Nobody scores that today, in their words

And the failure mode that hides from every dashboard: a fill_diagonal-style masking bug that excludes gold pairs from the candidate pool turns a retrieval eval into a constant-zero or constant-pass exercise. They shipped one. The test costs nothing: if every query scores zero hits, suspect your mask, not your model.

02 · FOUR EVAL PATTERNS, EACH PAID FOR

The second contribution (#1118) is four patterns from building a verified-memory system. We keep them in their original order because the order is the story: distribution, leakage, calibration, variance.

PatternTheir measurementThe recipe
benchmark scores do not transfer to your traffic bakeoff of bge-m3 vs Qwen3-embedding-0.6B on their own corpus (agent session memory, mixed zh/en): the MTEB-favored model gained +0.46pt recall@5 overall, under their transfer threshold, while the incumbent won recall@1 and the Chinese slices by +2.75pt re-rank embedders on a frozen sample of your own traffic before switching; public benchmarks are priors, not verdicts
group by qid or your eval leaks cross-item contamination (train items from the same task family appearing in eval retrieval context) silently inflated their early numbers every eval split is grouped by question or task id, never by row
three-valued retrieval beats binary their judge emits an unclear state with a confidence: Brier 0.069 with honest abstention, and the 2/27 abstentions in their latest public measurement were exactly the boundary rows let the retriever say don't know; score abstention as part of calibration (ECE), not as failure
to measure system drift, cache the inputs same-day score swings of 0.93 → 0.30 → 0.90 looked like model drift; replaying cached inputs through the judge K times showed 0.0pp judge-side variance, all of it in the generated side two-stage variance attribution, replay-cached-input vs regenerate-then-judge, before blaming retrieval or judging

03 · PREREGISTER, LABEL, BASELINE, PAIR

The third theme (#1138, recipe 2) is the discipline around the run itself, stated in four rules:

The ruleThe scar behind it
freeze criteria before running: metric, threshold, stop-loss, and the exact verifiable artifact each claim points to. Criteria may only move stricter mid-stream; loosening needs a fresh preregistration every rule here is a re-learned lesson from a run that was allowed to move its own goalposts
label every claim measured, inferred, or unverifiable. Inferred claims carry an upgrade path the most useful eval artifact they ever published was a self-report reading "2 of 4 gates FAIL on recompute". The visible FAILs are what made the green credible afterwards
compute the majority-class baseline before claiming accuracy a 0.87 binary classifier that a 0.78 always-yes baseline nearly matched, n=60, confidence intervals overlapping. Below n=100 the CI belongs next to every headline number, or the number is not a claim, it is a mood
paired designs beat unpaired by an order of magnitude of information same frames, two checkpoints, McNemar. Their unpaired reading overstated an effect by roughly 4× versus the paired one

The affinity is not an accident. This repo gates its own releases the same way: bars written before the work, FAILs kept visible in the notes, and a CI job that re-checks every claim in a release section against the tree it describes. The rc.80 gates post is the same philosophy with our numbers in it: three gates passed, three failed, one split, all published. When Chunxiao writes that visible FAILs are what made the green credible, that is the operating manual we already live by, arriving independently from a different domain. That convergence is worth more than either set of numbers alone.

04 · THE QUESTION THEY ASKED BACK, ANSWERED HONESTLY

Their price for the recipes was not a star. It was this: an honest issue pointing at where the index is thin against these patterns, especially session-memory slicing and failure-turn retention. We checked the registry rather than answering from memory. The answer, posted in full in #1138:

0
agent-trajectory corpora at the hub, of 99 live
0
session or trajectory entries in the backlog, of 132
file / section
our chunk unit, not the semantic topic their recall@1 needs

Three thin spots, stated plainly. First, the corpus they described, real multi-session agent trajectories with timestamps and failure turns kept verbatim, does not exist at our hub: zero of the 99 live manifests and none of the other 131 backlog entries are session or trajectory data. Nobody using our reference corpora can score what they score. Second, our indexer chunks at file and section boundaries and keeps a locator back to the source span; for session logs that is the wrong shape, and we have no semantic-topic slicer to offer instead. Third, our memory features store curated facts an agent chose to keep, with BM25 recall by default; there is no native tried-and-failed marker and no retention of turns that went nowhere.

What we did about it: the wanted corpus now has a named slot in the hub backlog (agent-session-trajectories, status planned, demand-anchored to their issue), and if Assay can donate sanitized multi-session trajectories with failure turns intact, the contribution path is open by pull request. It would be a first-of-kind corpus at the hub. Their offer to recompute our benchmarks independently stands unaccepted for now, and it should not: an independent recompute of any number we publish is always welcome.

What is claimed here, and by whom. Claimed by our guest, from their production system, with artifacts public in their repository: every measurement in §01 through §03, quoted from issues #1118 and #1138 with light editing and linked at each section. Claimed by us: the registry counts in §04 (98 manifests, backlog 132 rows after the wishlist entry, 99 live, zero matching a session, trajectory, chat or conversation query), verified 2026-10-07 by running the hub's own validator against the corpus-hub branch; the description of our chunking and memory features, which is a description of the code as it ships. Not claimed: that we have independently reproduced any of their numbers, that our engine today serves their workload well, or that the planned corpus exists in any form. It does not. That is the point of the backlog entry.

If you have production scars of your own: the same invitation stands. Open an issue with what it cost you and what you measured, and if it survives our check against reality, it gets the same treatment, your name on it. The engine itself is one binary with a quickstart, and the agent-memory recipe is the closest thing we ship today to what §01 describes wanting.

Guest material verbatim in issues #1118 and #1138 · their repository: nautilus-compass · the honest-gap comment: #1138 comment · backlog entry PR #1210 · related: the rc.80 gates · the Corpus Hub · the agent-memory recipe.