The proof-of-citation benchmark.

Merit's verifier scored on a fixed, hand-labeled gold set — a small, reproducible, falsifiable baseline, not a self-reported number. Every figure here regenerates from a fresh checkout with npm run judge-eval.

Fixed gold set

gold-set pairs
adversarial attacks held
live hard cases harvested
Precision / recall — loading…
The gold set mixes SUPPORTED pairs a correct verifier must pay and REFUSED ones it must hold — off-topic, contradiction, a fabricated figure, and the on-topic-but-contradictory trap that only an adversarial judge (not a similarity filter) catches. It is deliberately small and public so anyone can falsify the claim.

Self-bootstrapping set — boundary cases harvested from live traffic

Every run logs the citations the verifier was least sure about as gold-set candidates, so the benchmark co-evolves with real adversarial traffic instead of staying a static snapshot. Newest first.

verdictclaim
loading…
Coming: published external suites — RAGTruth, FaithBench, FACTS Grounding — scored by the same engine (npm run bench-judge). Drop the datasets under benchmark/ and the harness reports balanced-accuracy / precision / recall / F1 with confidence intervals. Until then this page reports only the fixed gold set, labeled with its size.
Want to move a number? Try to get a lie paid → Every attempt that survives becomes a new gold-set case.