Health Optimization Bench

Incretin Therapeutics: Evidence Synthesis

micro bench v1 · 89 released tasks · updated August 16, 2026

The suite covers the incretin therapeutic class: pivotal trial results, cardiovascular and renal outcome evidence, label facts, and guideline positions for GLP-1 receptor agonists and dual agonists, anchored to readouts from 2024 through 2026. Tasks demand specific quantities and study design details, not summaries, and each carries one penalized misstatement criterion scored without compensation.

Ranking

#modelscore95% CIsafety failsunanimity
1Anthropic logoClaude Fable 5 Anthropic83.879.188.110/8991%
2xAI logoGrok 4.6 xAI81.377.085.312/8988%
3Anthropic logoClaude Opus 5 Anthropic78.372.983.414/8989%
4OpenAI logoGPT-5.6 Sol (max) OpenAI77.973.382.213/8989%
5Moonshot AI logoKimi K3 Moonshot AI77.873.082.414/8989%
6OpenAI logoGPT-5.6 Sol (high) OpenAI77.773.182.013/8989%
7Meta logoMuse Spark Meta68.462.873.919/8986%
8Google logoGemini 3.6 Google57.951.264.319/8981%
9TMInkling Thinking Machines53.947.760.123/8982%
10Anthropic logoClaude Sonnet 5 Anthropic50.144.356.120/8982%
11MiniMax logoMiniMax M3 MiniMax35.529.641.634/8983%
12Microsoft AI logoMAI Thinking Microsoft AI33.026.839.327/8985%
13Zhipu logoGLM 5.2 Zhipu29.223.934.933/8983%
14Mistral logoMistral Medium 3.5 Mistral14.410.618.547/8989%
15NVIDIA logoNemotron 3.5 Lightning NVIDIA7.84.711.348/8994%

Consensus rubric credit, one answer per task, tools disabled. Safety fails count tasks with a consensus-met penalty criterion. Unanimity is the share of criterion judgments on which all three grader families agreed before adjudication.

Task difficulty

hardest, cross-model mean

  • redefine1 projection audit13%
  • placebo weight trajectories13%
  • pivotal dc ae extract18%
  • society guideline grades18%
  • retatrutide phase2 status22%

easiest, cross-model mean

  • flow early termination91%
  • incretin label indications90%
  • select mace nnt a286%
  • dcae pivotal obesity trials85%
  • select mace nnt a484%

Construction and grading follow the bench-wide protocol: five-family authoring, two-family confirmation before any removal, blind author-excluded panel grading with escalation, and a confidential holdout. Details are on the methodology page.