Health Optimization Bench

Methodology

The benchmark is built on one premise: no single model family can be trusted to write, check, or score its own examination. Authoring, auditing, and grading are therefore distributed across independent frontier families at every stage, and the numbers below are measured from the pipeline itself rather than asserted.

Construction pipeline

AuthoringTasks are written one at a time, each by a single frontier model family, five families in total, stratified so that no family's style correlates with a topic. The v1 evidence suite began at 120 tasks.
Structural checksDeterministic code, never a model: schema validity, five to seven binary criteria, exactly one penalty criterion, and a cited source for every criterion.
Gold auditA family that did not write the task reviews it for material error against its cited sources.
Citation auditA second non-author family checks each criterion against the passage it cites.
Confirmation ruleA material verdict removes a task only when a second family independently confirms it. Verdicts flagged by one family and cleared by another are retained and recorded.
Difficulty bandingTwo-sided and preregistered. Saturated tasks are excluded, and so are tasks that no probe model can pass, since universal failure usually signals ambiguity rather than difficulty.
HoldoutA reserved subset of tasks never transits an evaluated vendor's infrastructure before that vendor's official run.
Release89 tasks released for the v1 evidence suite; 17 held out.

Measured verification outcomes

Residual material-error rate8%tasks removed after two-family confirmation, 10 of 120
Single-auditor false positives35 rescuedflagged by one family, cleared by a second on independent review
Cross-family grading unanimity87%criterion judgments on which all three grader families agree
Adjudication rate~half of tasksat least one criterion escalated to a fourth family
Task difficulty range13–91%cross-model mean per task; no ceiling and no floor cluster

Grading protocol

Candidate models answer closed book, with tools disabled, one answer per task. Grading is blind: the grader sees the task, the candidate answer, and the criterion text only. It never sees point values, the reference answer, or another grader's judgment. Three passes are distributed across three different families, and for any task the authoring family is excluded from the panel and substituted. Unanimous verdicts stand. Any split escalates the criterion to a fourth, uninvolved family, whose vote joins the majority decision. Criterion order is shuffled per pass, and prompt versions are recorded with every run.

Safety scoring

Every task carries one penalty criterion describing a specific, plausible misstatement of the evidence. It is scored without compensation: credit earned elsewhere never offsets it, and the leaderboard reports these failures as a separate column. A model can write a moderately scoring answer while misstating trial results on half the tasks, and the two behaviors are reported apart so that neither hides the other.

Comparability

Scores are comparable only within an identical task set, grading panel, and pass count. Run identity is deterministic, so re-running the same configuration on the same day reproduces the same run record. Confidence intervals are nonparametric bootstrap intervals over tasks with a fixed seed.

Limitations

Single-sample scoring measures typical rather than best-of-k performance; a pass@k extension on a preregistered subset is planned. Tasks anchored to 2024 through 2026 evidence confound recency with reasoning for models with earlier training cutoffs, which we treat as part of what the bench measures and report per task. The clinical decision track, covering management, dosing, and refusal scenarios, is withheld from this release while it completes licensed-clinician review, on the principle that clinical validity is certified by physicians rather than by model consensus.