XERJ · BLOG · THE ENGINEERING LOG

MEASURED,
WRITTEN DOWN.

Results with the losses left in. Newest first.

This is where the measurements get written up as they happen — the wins, the losses, and the bill. Every post is dated, cites the runs it came from, and leaves the losing rows in the tables, because a number you cannot trace is marketing. The entry below puts TypeSafe AI's Jev judge model over 1,271 judged BEIR queries against a XERJ first stage, then turns the wire around and answers it from a local node: no tokens, no egress, every figure reproduced from the artifacts it names.

  1. · BENCHMARKS · JEV

    Does XERJ beat JEV? On the bill, outright. On FiQA, no.

    Does XERJ beat JEV? On the bill, outright — the FiQA rerank run billed $0.6375 and 15,178,758 judge tokens, while our local node answers the same wire with zero. On relevance the hosted judge keeps SciFact (0.7410 vs 0.6993) and FiQA (0.3638 vs 0.2382, +0.126); our zero-API-token hybrid keeps NFCorpus (0.3448 vs 0.3312). And the unmodified pip jev-reranker passed its own gate against a XERJ node. Every win, every loss, with the receipt.

    READ THE POST →

More measured ground lives in the benchmarks hub: /benchmarks — the ES conformance board, BEIR baselines, and the Jev tables, every number traced to a run.