- Input
- Identifiers and FHIR search functions
- Output
- Requested value or identifier
- Measure
- Rule-based query SR
Medical agent benchmark analysis
Read the action behind the answer.
Independent analysis of MedAgentBench and AgentClinic: task success, simulated dialogue, query/action differences and the limits of medical agent scores.
Reading a record. Proposing an action.
MedAgentBench ↗arXiv v2, February 2025; 150 queries, pass@1, eight-round baseline; temperature zero for these selected models. Each model has a query score and a separate action score. The connecting line pairs its two measurements.
Historical selected rows. Table 3 does not provide uncertainty intervals. Keep this separate from query scores. No deployment reliability or current model ranking is implied. Table 3. [1]
MedAgentBench
- Resolve the task
Read instruction and hospital context before choosing a FHIR function. [1]
- Retrieve
GET requests return raw server responses to the agent. [1]
- Propose an action
POST payloads receive local checks and a simulated success response; baseline writes are not executed. [1]
- Finish and grade
Compare query answers with reference solutions; apply authored checks to action payloads. [1]
AgentClinic
- Distribute case information
Different case fields are exposed to different simulated roles. [4]
- Gather evidence
The doctor asks the patient questions or requests measurements. [4]
- Commit to a diagnosis
The doctor concludes before the interaction budget is exhausted. [4]
- Judge the conclusion
The moderator compares the conclusion with the reference diagnosis. [4][6]
Published results
Source record ↗Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.
Paper-reported results / selected rows
Historical AgentClinic-MedQA comparison
2026 journal article, Figure 2; GPT-4 simulated patient and measurement agents, 20-interaction cap.
Published measurements only. The ± quantities are reproduced as reported, not relabeled as confidence intervals.
Source: Comparison of models; Figure 2 [4]
An original analytical tool
Which part of an agent does the benchmark exercise?
Filter task families to compare the information available, expected behavior and point where grading stops. These annotations are our analysis of published protocols.
7 of 7 evidence entries shown
- Input
- Time-bounded observations
- Output
- Computed task answer
- Measure
- Rule-based query SR
- Input
- New observation and patient identifier
- Output
- POST payload
- Measure
- Action SR
- Input
- Task conditions, codes and retrieved context
- Output
- Medication/test/referral payload
- Measure
- Action SR
- Input
- Partial information distributed among agents
- Output
- Final diagnosis after dialogue
- Measure
- Diagnostic accuracy
- Input
- Doctor requests to the measurement agent
- Output
- Returned finding, then diagnosis
- Measure
- Diagnostic accuracy
- Input
- NEJM-derived image and case interaction
- Output
- Final diagnosis
- Measure
- Diagnostic accuracy
Coverage rows describe evaluation scope. They are not measured performance or proof that a capability transfers to a deployed workflow. [1][4]
The benchmark in detail
All dossiers →300-task paper; arXiv v2
MedAgentBench ↗
A successful read and a successful action are different claims.
2026 journal article; expanded inventory from v5 appendix
AgentClinic ↗
The simulated patient is part of the test instrument.
What we examine
Explore two different forms of agency: FHIR record tasks in MedAgentBench and diagnostic dialogue in AgentClinic. Our original analysis tracks what an agent receives, what it must do and what the grader actually checks. The published results retain their historical model configurations and benchmark versions. Use the coverage explorer to locate a capability, then inspect why successful retrieval, valid proposed actions and correct simulated diagnoses support different conclusions.
- Inspect the task
- Trace the actual input, output and evaluation setting for each named benchmark.
- Read the evidence
- Explore cited cohort counts and selected historical results with their measurement boundaries.
- Make the inference explicit
- Separate the authors’ observations from our analytical interpretation and proposed evaluation questions.
Analysis & interpretation
All analyses →What does MedAgentBench success actually measure?
Read MedAgentBench’s query and action results through its 300-task protocol, parser and proposed-write grading boundary.
Why is the simulated patient part of an AgentClinic result?
Separate the tested doctor model, simulated patient, measurement agent and diagnosis moderator when interpreting AgentClinic.
How should MedAgentBench and AgentClinic be compared?
Build a capability comparison between record actions and diagnostic dialogue without averaging non-equivalent benchmark scores.
Questions, answered
Are these the official benchmark websites?
No. This is an independent analysis of MedAgentBench and AgentClinic. The original authors, papers and repositories are cited throughout.
Does MedAgentBench action success mean an EHR write was executed?
In the analyzed 300-task paper baseline, proposed POST payloads are checked locally and not executed against the server. Action success must be read within that grading boundary.
Why keep query and action results separate?
They grade different task subgroups and can rank the same models differently. Their equal mixture in the benchmark is not a universal deployment workload.
Does AgentClinic evaluate real patient conversations?
It evaluates simulated interactions grounded in selected case sources. Clinical record-derived scenarios do not turn generated dialogue into observed patient behavior.
Prepare a comparison worksheet
Working tool / saved on this device
Prepare a reproducible benchmark reading
Use this secondary worksheet to record the version, task and evidence boundary of the benchmark you are considering. Checked items indicate documented decisions, not measured performance.
Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.
Download the evidence ↗