Healthcare AI governance training and certification. One day online, Oct 19. | Register now →

Webinar

Agentic Testing of Medical AI: HealthAdminBench, PhysicianBench, and CliniCARE-Bench in MedHELM

Watch live
October 28, 2026 @ 2:00 PM ET
Medical AI is moving from answering questions to doing work: processing prior authorizations, working a specialist consultation across dozens of tool calls, auditing a chart against a clinical standard. A single-turn benchmark cannot show if an agent completed the workflow it was given, effectively planning and using tools. Current results show the gap is large: on HealthAdminBench, agents complete 83% of individual subtasks but only 36% of workflows end to end. On PhysicianBench, the best of 13 agents succeeds on 46% of tasks. on CliniCARE-Bench, every one of 16 systems tested committed to an answer even when the patient record doesn’t support one.”This session walks through the three agentic benchmarks now integrated into the open-source MedHELM framework. They all score what the agent actually did rather than what it said it would do:
  • HealthAdminBench (Stanford, Shah Lab): 135 multi-step workflows across prior authorization, appeals and denials, and DME ordering, decomposed into 1,698 verifiable subtasks and executed by computer use agents in a browser.
  • PhysicianBench (Stanford, HealthRex lab): 100 long-horizon tasks adapted from real primary-care-to-subspecialty consultations across 21 specialties, averaging 27 tool calls per task.
  • CliniCARE-Bench (Scale AI): 25 clinician-authored audit scenarios instantiated as 750 patient cases over MIMIC-IV data, with two indeterminate verdict classes that make principled abstention part of the score.
You’ll learn:
  • How the extended MedHELM harness evaluates agents that plan, call tools, and work across many turns.
  • How to run the benchmarks yourself: install, configure your model, and execute a workflow, and populate a leaderboard.
  • Why one open-source platform for accuracy, agentic reliability, safety, and bias testing beats a patchwork of tools, and how the same suites run as a CI/CD gate with Gatekeeper and as a production monitor with Guardian.

Webinar Sign Up


About the speaker
Alin Blidisel
Engineering Technical Lead at Pacific AI

Alin Blidisel is an AI Infrastructure & Evaluation Lead at Pacific AI focused on large-scale AI infrastructure and data systems. He has built scalable cloud and big data architectures across AWS, Google Cloud, and Microsoft Azure. His work spans machine learning infrastructure, distributed data processing, and production AI systems for healthcare, with a recent focus on integrating new frontier models into evaluations and orchestrating test scenarios.