—
gold-set pairs
Merit's verifier scored on a fixed, hand-labeled gold set — a small, reproducible, falsifiable baseline, not a self-reported number. Every figure here regenerates from a fresh checkout with npm run judge-eval.
Every run logs the citations the verifier was least sure about as gold-set candidates, so the benchmark co-evolves with real adversarial traffic instead of staying a static snapshot. Newest first.
| verdict | claim |
|---|---|
| loading… | |
npm run bench-judge). Drop the datasets under benchmark/ and the harness reports balanced-accuracy / precision / recall / F1 with confidence intervals. Until then this page reports only the fixed gold set, labeled with its size.