Methodology
The benchmark is built on one premise: no single model family can be trusted to write, check, or score its own examination. Authoring, auditing, and grading are therefore distributed across independent frontier families at every stage, and the numbers below are measured from the pipeline itself rather than asserted.
Construction pipeline
| Authoring | Tasks are written one at a time, each by a single frontier model family, five families in total, stratified so that no family's style correlates with a topic. The v1 evidence suite began at 120 tasks. |
| Structural checks | Deterministic code, never a model: schema validity, five to seven binary criteria, exactly one penalty criterion, and a cited source for every criterion. |
| Gold audit | A family that did not write the task reviews it for material error against its cited sources. |
| Citation audit | A second non-author family checks each criterion against the passage it cites. |
| Confirmation rule | A material verdict removes a task only when a second family independently confirms it. Verdicts flagged by one family and cleared by another are retained and recorded. |
| Difficulty banding | Two-sided and preregistered. Saturated tasks are excluded, and so are tasks that no probe model can pass, since universal failure usually signals ambiguity rather than difficulty. |
| Holdout | A reserved subset of tasks never transits an evaluated vendor's infrastructure before that vendor's official run. |
| Release | 89 tasks released for the v1 evidence suite; 17 held out. |
Measured verification outcomes
| Residual material-error rate | 8% | tasks removed after two-family confirmation, 10 of 120 |
| Single-auditor false positives | 35 rescued | flagged by one family, cleared by a second on independent review |
| Cross-family grading unanimity | 87% | criterion judgments on which all three grader families agree |
| Adjudication rate | ~half of tasks | at least one criterion escalated to a fourth family |
| Task difficulty range | 13–91% | cross-model mean per task; no ceiling and no floor cluster |
Grading protocol
Candidate models answer closed book, with tools disabled, one answer per task. Grading is blind: the grader sees the task, the candidate answer, and the criterion text only. It never sees point values, the reference answer, or another grader's judgment. Three passes are distributed across three different families, and for any task the authoring family is excluded from the panel and substituted. Unanimous verdicts stand. Any split escalates the criterion to a fourth, uninvolved family, whose vote joins the majority decision. Criterion order is shuffled per pass, and prompt versions are recorded with every run.
Safety scoring
Every task carries one penalty criterion describing a specific, plausible misstatement of the evidence. It is scored without compensation: credit earned elsewhere never offsets it, and the leaderboard reports these failures as a separate column. A model can write a moderately scoring answer while misstating trial results on half the tasks, and the two behaviors are reported apart so that neither hides the other.
Comparability
Scores are comparable only within an identical task set, grading panel, and pass count. Run identity is deterministic, so re-running the same configuration on the same day reproduces the same run record. Confidence intervals are nonparametric bootstrap intervals over tasks with a fixed seed.
Limitations
Single-sample scoring measures typical rather than best-of-k performance; a pass@k extension on a preregistered subset is planned. Tasks anchored to 2024 through 2026 evidence confound recency with reasoning for models with earlier training cutoffs, which we treat as part of what the bench measures and report per task. The clinical decision track, covering management, dosing, and refusal scenarios, is withheld from this release while it completes licensed-clinician review, on the principle that clinical validity is certified by physicians rather than by model consensus.