Health Optimization Bench

A rubric-graded benchmark of frontier language models on current clinical evidence.

Leaderboard

v1 release set, 89 tasks. Updated August 16, 2026.

83.8
81.3
78.3
77.9
77.8
68.4
57.9
53.9
50.1
35.5
33.0
29.2
14.4
7.8
Anthropic logo
xAI logo
Anthropic logo
OpenAI logo
Moonshot AI logo
Meta logo
Google logo
TM
Anthropic logo
MiniMax logo
Microsoft AI logo
Zhipu logo
Mistral logo
NVIDIA logo
Claude Fable 5
Grok 4.6
Claude Opus 5
GPT-5.6 Sol (max)
Kimi K3
Muse Spark
Gemini 3.6
Inkling
Claude Sonnet 5
MiniMax M3
MAI Thinking
GLM 5.2
Mistral Medium 3.5
Nemotron 3.5 Lightning

Consensus rubric credit on the release set, one answer per task, tools disabled. 95 percent bootstrap confidence intervals over tasks accompany every score on the full ranking. GPT-5.6 Sol (high) is completing its run and will be added when grading finishes.

Health Optimization Bench evaluates language models on hard, freshness-dependent questions in preventive and optimization medicine. Every task is written against a primary source, audited by model families that did not author it, and scored blind by a panel of three independent families. The task's authoring family never grades it.

15
models evaluated
12
labs represented
89
released tasks
120
tasks authored
17
task confidential holdout
5
authoring model families

Micro benches

The bench is assembled from focused micro benches, each covering one area of preventive and optimization medicine. The incretin therapeutics evidence suite is the first with a released model ranking. The remaining suites are authored and are moving through the same cross-family verification pipeline; clinical decision tasks are additionally reviewed by licensed clinicians before release.

Data sample

A 30 task sample, drawn from each micro bench, is available through the Arcophos research data platform, with rubric criteria and reference responses included.

View the sample

Paper

The methodology paper is in preparation and will accompany the archived v1 release, together with per-task verification records and the grading protocol. For early access, licensing, or evaluation of a private model, contact Arcophos.