Medical AI is moving from answering questions to doing work: processing prior authorizations, working a specialist consultation across dozens of tool calls, auditing a chart against a clinical standard. A single-turn benchmark cannot show if an agent completed the workflow it was given, effectively planning and using tools. Current results show the gap is large: on HealthAdminBench, agents complete 83% of individual subtasks but only 36% of workflows end to end. On PhysicianBench, the best of 13 agents succeeds on 46% of tasks. on CliniCARE-Bench, every one of 16 systems tested committed to an answer even when the patient record doesn’t support one.”
This session walks through the three agentic benchmarks now integrated into the open-source MedHELM framework. They all score what the agent actually did rather than what it said it would do:
- HealthAdminBench (Stanford, Shah Lab): 135 multi-step workflows across prior authorization, appeals and denials, and DME ordering, decomposed into 1,698 verifiable subtasks and executed by computer use agents in a browser.
- PhysicianBench (Stanford, HealthRex lab): 100 long-horizon tasks adapted from real primary-care-to-subspecialty consultations across 21 specialties, averaging 27 tool calls per task.
- CliniCARE-Bench (Scale AI): 25 clinician-authored audit scenarios instantiated as 750 patient cases over MIMIC-IV data, with two indeterminate verdict classes that make principled abstention part of the score.
You’ll learn:
- How the extended MedHELM harness evaluates agents that plan, call tools, and work across many turns.
- How to run the benchmarks yourself: install, configure your model, and execute a workflow, and populate a leaderboard.
- Why one open-source platform for accuracy, agentic reliability, safety, and bias testing beats a patchwork of tools, and how the same suites run as a CI/CD gate with Gatekeeper and as a production monitor with Guardian.