ROBERT DURANIV
RDIV / STRATEGY / RESEARCHLIVE ARCHIVE

R/IV / VERIFICATION BENCHMARK

MEASURE THE
QUALITY OF
VERIFICATION.

The current benchmark is a synthetic calibration harness for testing contracts, failure handling, and appropriate abstention. It is not a comparative leaderboard.

SYNTHETIC CALIBRATION · REPRODUCIBLE METHOD · NO CLAIMED SUPERIORITY

[ EVALUATION TASKS ]

TEST THE RECORD.
TEST THE LIMITS.

Each case fixes its inputs, expected evidence boundary, scoring rule, and failure condition before execution.

01 / TASK

CITATION EXISTENCE

Can the cited record be resolved and identified?

02 / TASK

CITATION ENTAILMENT

Does the source support the bounded claim attributed to it?

03 / TASK

UNSUPPORTED CLAIMS

Which material statements outrun admitted evidence?

04 / TASK

CONTRADICTIONS

Which evidence objects materially conflict?

05 / TASK

PROVENANCE

Can each material finding be traced to its admitted record?

06 / TASK

ACTION VERIFICATION

Does independent evidence establish that an action occurred?

07 / TASK

ABSTENTION

Does the system stop when the record is insufficient?

[ REPRODUCIBILITY MANIFEST ]

VERSION THE TEST.
PRESERVE THE LIMITS.

Benchmark version, task IDs, lawful-source policy, expected behavior, scoring rules, system release, run timestamp, provider/model version, false-positive analysis, false-negative analysis, and abstention behavior are recorded in the machine-readable manifest. The latest run remains null until a valid execution occurs.

[ PUBLICATION RULE ]

CALIBRATION FIRST.
COMPARISON LATER.

Comparative scores will remain unpublished until a frozen, lawful real-world corpus exists and the full run is reproducible. No private holdout material, fabricated result, competitor ranking, or unsupported accuracy claim is published here.