Skip to main content
Each battery ships its labelled rows, and a receipt for each engine it was measured with (hunch test --model <engine> --receipt). This page is those receipts side by side. A row is one engine answering one question on the battery’s own rows; accuracy and AUROC are against the battery’s gold. An engine’s number here is on this battery’s data: measure it on yours before choosing. act bars are set for the spec’s own engine; another engine’s probabilities mean something else, so its acting numbers aren’t comparable and are left out.

agent-commands

Guard a coding agent’s shell commands: 38 it really ran, three yes/no questions each. The quickstart; about $0.001 to run.

rag-answers

Check a RAG answer against its sources: does it say anything the passages don’t support? 900 real answers marked by annotators, one yes/no question; AUROC 0.941; about $0.03 to run.

tickets

Route 40 made-up support tickets: which team, is it urgent, how frustrated (a choice, a yes/no and a score). About $0.001 to run.