Medical agent benchmark analysis

Read the action behind the answer.

Independent analysis of MedAgentBench and AgentClinic: task success, simulated dialogue, query/action differences and the limits of medical agent scores.

Independent analysis by Arcophos · updated

Reading a record. Proposing an action.

MedAgentBench ↗

arXiv v2, February 2025; 150 queries, pass@1, eight-round baseline; temperature zero for these selected models. Each model has a query score and a separate action score. The connecting line pairs its two measurements.

Task success rate, 0–100% · 150 queries and 150 proposed actions [1]
QueryProposed action
Model
050100%
QueryAction
Claude 3.5 Sonnet v2
85.3354.00
GPT-4o
72.0056.00
DeepSeek-V3
70.6754.67
Gemini-1.5 Pro
52.6771.33

Historical selected rows. Table 3 does not provide uncertainty intervals. Keep this separate from query scores. No deployment reliability or current model ranking is implied. Table 3. [1]

MedAgentBench

  1. Resolve the task

    Read instruction and hospital context before choosing a FHIR function. [1]

  2. Retrieve

    GET requests return raw server responses to the agent. [1]

  3. Propose an action

    POST payloads receive local checks and a simulated success response; baseline writes are not executed. [1]

  4. Finish and grade

    Compare query answers with reference solutions; apply authored checks to action payloads. [1]

AgentClinic

  1. Distribute case information

    Different case fields are exposed to different simulated roles. [4]

  2. Gather evidence

    The doctor asks the patient questions or requests measurements. [4]

  3. Commit to a diagnosis

    The doctor concludes before the interaction budget is exhausted. [4]

  4. Judge the conclusion

    The moderator compares the conclusion with the reference diagnosis. [4][6]

Published results

Source record ↗

Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.

Paper-reported results / selected rows

Historical AgentClinic-MedQA comparison

2026 journal article, Figure 2; GPT-4 simulated patient and measurement agents, 20-interaction cap.

Diagnostic accuracy · %
050100
Reported
Claude-3.5Paper reports ±3.3 alongside the mean.
62.1%
GPT-4Paper reports ±3.3 alongside the mean.
51.6%
GPT-4oPaper reports ±3.4 alongside the mean.
34.2%

Published measurements only. The ± quantities are reproduced as reported, not relabeled as confidence intervals.

Source: Comparison of models; Figure 2 [4]

An original analytical tool

Which part of an agent does the benchmark exercise?

Evidence explorer

Filter task families to compare the information available, expected behavior and point where grading stops. These annotations are our analysis of published protocols.

7 of 7 evidence entries shown

Record retrieval

Find a patient or result

Read dossier ↗
Input
Identifiers and FHIR search functions
Output
Requested value or identifier
Measure
Rule-based query SR
Interpretation boundary

Depends on the parser and reference answer.

[1]
Record retrieval

Aggregate recent measurements

Read dossier ↗
Input
Time-bounded observations
Output
Computed task answer
Measure
Rule-based query SR
Interpretation boundary

Tests the specified calculation and time filter.

[1]
Proposed actions

Record a measurement

Read dossier ↗
Input
New observation and patient identifier
Output
POST payload
Measure
Action SR
Interpretation boundary

Baseline validates payload, without executing a server write.

[1]
Proposed actions

Prepare an order

Read dossier ↗
Input
Task conditions, codes and retrieved context
Output
Medication/test/referral payload
Measure
Action SR
Interpretation boundary

Does not verify downstream care or persisted server state.

[1]
Diagnostic dialogue

Elicit a diagnostic history

Read dossier ↗
Input
Partial information distributed among agents
Output
Final diagnosis after dialogue
Measure
Diagnostic accuracy
Interpretation boundary

Patient-model responses influence the test.

[4]
Diagnostic dialogue

Request a diagnostic measurement

Read dossier ↗
Input
Doctor requests to the measurement agent
Output
Returned finding, then diagnosis
Measure
Diagnostic accuracy
Interpretation boundary

Information request and answer quality are coupled.

[4]
Multimodal dialogue

Interpret a case image

Read dossier ↗
Input
NEJM-derived image and case interaction
Output
Final diagnosis
Measure
Diagnostic accuracy
Interpretation boundary

Curated case challenges are not routine imaging prevalence.

[4][5]

Coverage rows describe evaluation scope. They are not measured performance or proof that a capability transfers to a deployed workflow. [1][4]

The benchmark in detail

All dossiers →
Dossier01

300-task paper; arXiv v2

MedAgentBench ↗

A successful read and a successful action are different claims.

UnitOne patient-specific task attempt.MeasureTask success rate
Dossier02

2026 journal article; expanded inventory from v5 appendix

AgentClinic ↗

The simulated patient is part of the test instrument.

UnitOne simulated diagnostic encounter.MeasureDiagnostic accuracy

What we examine

Explore two different forms of agency: FHIR record tasks in MedAgentBench and diagnostic dialogue in AgentClinic. Our original analysis tracks what an agent receives, what it must do and what the grader actually checks. The published results retain their historical model configurations and benchmark versions. Use the coverage explorer to locate a capability, then inspect why successful retrieval, valid proposed actions and correct simulated diagnoses support different conclusions.

Inspect the task
Trace the actual input, output and evaluation setting for each named benchmark.
Read the evidence
Explore cited cohort counts and selected historical results with their measurement boundaries.
Make the inference explicit
Separate the authors’ observations from our analytical interpretation and proposed evaluation questions.

Analysis & interpretation

All analyses →

Questions, answered

Are these the official benchmark websites?

No. This is an independent analysis of MedAgentBench and AgentClinic. The original authors, papers and repositories are cited throughout.

Does MedAgentBench action success mean an EHR write was executed?

In the analyzed 300-task paper baseline, proposed POST payloads are checked locally and not executed against the server. Action success must be read within that grading boundary.

Why keep query and action results separate?

They grade different task subgroups and can rank the same models differently. Their equal mixture in the benchmark is not a universal deployment workload.

Does AgentClinic evaluate real patient conversations?

It evaluates simulated interactions grounded in selected case sources. Clinical record-derived scenarios do not turn generated dialogue into observed patient behavior.

Prepare a comparison worksheet

Working tool / saved on this device

Prepare a reproducible benchmark reading

Interactive worksheet

Use this secondary worksheet to record the version, task and evidence boundary of the benchmark you are considering. Checked items indicate documented decisions, not measured performance.

Identify the experiment

Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.

Download the evidence ↗