CITATION EXISTENCE
Can the cited record be resolved and identified?
R/IV / VERIFICATION BENCHMARK
The current benchmark is a synthetic calibration harness for testing contracts, failure handling, and appropriate abstention. It is not a comparative leaderboard.
SYNTHETIC CALIBRATION · REPRODUCIBLE METHOD · NO CLAIMED SUPERIORITY[ EVALUATION TASKS ]
Each case fixes its inputs, expected evidence boundary, scoring rule, and failure condition before execution.
Can the cited record be resolved and identified?
Does the source support the bounded claim attributed to it?
Which material statements outrun admitted evidence?
Which evidence objects materially conflict?
Can each material finding be traced to its admitted record?
Does independent evidence establish that an action occurred?
Does the system stop when the record is insufficient?
[ REPRODUCIBILITY MANIFEST ]
Benchmark version, task IDs, lawful-source policy, expected behavior, scoring rules, system release, run timestamp, provider/model version, false-positive analysis, false-negative analysis, and abstention behavior are recorded in the machine-readable manifest. The latest run remains null until a valid execution occurs.
[ PUBLICATION RULE ]
Comparative scores will remain unpublished until a frozen, lawful real-world corpus exists and the full run is reproducible. No private holdout material, fabricated result, competitor ranking, or unsupported accuracy claim is published here.