Results with the losses left in. Newest first.
This is where the measurements get written up as they happen — the wins, the losses, and the bill. Every post is dated, cites the runs it came from, and leaves the losing rows in the tables, because a number you cannot trace is marketing. The entry below puts TypeSafe AI's Jev judge model over 1,271 judged BEIR queries against a XERJ first stage, then turns the wire around and answers it from a local node: no tokens, no egress, every figure reproduced from the artifacts it names.
More measured ground lives in the benchmarks hub: /benchmarks — the ES conformance board, BEIR baselines, and the Jev tables, every number traced to a run.