What Adam Is Reading
The Benchmark Says Gemini. The Clinic Says Wait.
A new Nature Medicine paper finds that general purpose frontier models beat the specialized clinical AI tools hospitals are buying. The finding is probably true. What it licenses you to do about it is smaller than it looks.
Single source review with critique · 8 sources · August 2026

UpToDate Expert AI declined to answer nineteen of every hundred questions it was asked. The three frontier models declined between one and three. I have consulted with people at both ends of that range. Only one of them ever worried me, and it was not the one who occasionally said no.

That number sits in a figure panel most readers will skip, in a paper whose title does the heavy lifting. General purpose large language models outperform specialized clinical AI tools on medical benchmarks. Vishwanath and colleagues at NYU Langone and UT Austin put GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 against OpenEvidence and UpToDate Expert AI, threw in Google Search AI Overview as a control nobody asked for, and had twelve blinded clinicians grade 1,800 responses. The frontier models won everything.

I believe the result. I want to talk about what it measures, because the gap between "scored higher when a doctor read it" and "should be the thing a doctor uses" is where the entire clinical AI field currently lives.


What they actually did

Three stages. Five hundred USMLE style multiple choice questions from MedQA. Five hundred free response prompts from HealthBench, graded by a panel of three LLM judges. And the new one, the Real Clinical Queries benchmark, built from 100 de-identified questions that physicians actually typed into NYU Langone's internal HIPAA compliant GPT instance during ordinary work. Twelve US clinicians rated every response on clinical correctness, completeness, safety, and clarity, blinded to which system produced it, three raters per response.

System MedQA HealthBench RCQ (1 to 4)
Gemini 3.1 Pro97.4%79.33.62
GPT-5.294.2%88.03.54
Claude Opus 4.690.2%77.03.52
Google AI Overviewnot testednot tested3.27
OpenEvidence89.6%62.63.24
UpToDate Expert AI88.4%61.33.17
Two clean tiers on the clinical query benchmark, with no significant separation inside either one. The commercial clinical tools land in the same tier as the free thing that appears above your Google results, which is the sentence the vendors will spend the next quarter explaining.
The finding nobody put in a headline. Harmful content did not differ across the six systems (Cochran's Q = 4.00, P = 0.55). Neither did hallucination (Q = 5.00, P = 0.42). Claude, which finished third in overall quality, had the highest rate of responses flagged as potentially harmful at 3.0 percent. Quality ranking and safety ranking are not the same list. Governance committees care about the second one.

Where this breaks on contact with clinical care
1
The questions were pre-selected for the winner
What actually happened

The RCQ benchmark is the paper's crown jewel and its deepest structural problem. Those 100 questions came from physicians querying a general purpose GPT deployment. Clinicians learn within about a week what a chatbot is good at and route their questions accordingly. The narrative, synthesis, and explain it to me questions go to the chatbot. The vancomycin dosing in CRRT question goes to UpToDate, where it never enters the sampling frame.

Why it matters

This is a benchmark of questions selected by users of the winning technology, scored on how well that technology answers them. It is the most realistic query set anyone has built, and it still has home field advantage baked into the sampling.

Real finding, narrower scope
2
The outcome is graded prose, not a treated patient
What actually happened

Every number in this paper describes text that a clinician read and scored. No patient was seen. No decision was changed. No outcome was measured. Twelve raters worked from a four point rubric, and the item level agreement was fair at best, with Krippendorff's alpha between 0.10 and 0.21. Collapse the scale to acceptable versus unacceptable and agreement improves substantially, which tells you the instrument reliably separates good from bad and struggles to separate good from very good.

Why it matters

The headline differences on the clinical query benchmark are 0.36 to 0.44 points on a four point scale. The paper is asking a noisy instrument to resolve a distinction finer than its demonstrated grain. The tier separation is robust. The ordering inside the top tier is not, and the authors are careful to say so.

Directionally sound, over-resolved
3
One turn, no context, no chart
What actually happened

Each query was submitted once and the response was graded as delivered. Real clinical question answering is iterative. You ask, you get something adjacent, you narrow, you check the citation, you decide. The dimension where models separated most was clarity (Kendall's W = 0.292). The dimension where they separated least was clinical correctness (W = 0.141).

Why it matters

The strongest and most reliable thing this paper demonstrates is that frontier models write better. That is not trivial. It is also not the same as knowing more medicine, and the paper's own numbers say so. OpenEvidence scored lowest on clarity at 2.84, and the authors read that as a communication weakness rather than a knowledge one.

Honest limitation, well disclosed
4
Unequal access to the contestants
What actually happened

Frontier models were queried through APIs at temperature zero with a fixed seed and search tools enabled. OpenEvidence and UpToDate have no public API, so they were queried by hand through browser interfaces, with hidden system prompts, unknown retrieval settings, and whatever output formatting the vendor chose that week.

Why it matters

The authors acknowledge this and it is not fixable from outside. It also means part of the measured gap is a comparison of a raw model against a productized wrapper, and wrappers get updated. The clinical tools were sampled twice, months apart, which is more diligence than most evaluations manage and still a snapshot of a moving object. By the time this appeared in print in June, GPT-5.2 and Gemini 3.1 were themselves historical.

Unavoidable, still a confound
5
Refusals were dropped, and refusal is clinical behavior
What actually happened

Thirty two question and model pairs were flagged as refusals and excluded, leaving 568 scored responses. UpToDate accounted for most of them at a 19 percent refusal rate, against 1 to 3 percent for the frontier models and 6 percent for Google.

Why it matters

Refusal is scored on a separate ledger from answer quality, which means a tool that declines to guess is penalized once and quietly rewarded in the other column, since its hardest cases never get graded. Neither extreme is the right answer. A tool that answers 99 percent of clinical questions is either extraordinary or insufficiently worried, and the paper cannot tell you which. Calibrated refusal is a feature I would pay for. It is also, at 19 percent, an unusable product.

Reported cleanly, interpreted thinly
6
The judges went to school with the contestants
What actually happened

MedQA and HealthBench were scored by a panel of three LLM judges, and the panel was GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. The same three systems being evaluated. HealthBench itself was built by OpenAI, and GPT-5.2 posted the highest score on it.

Why it matters

The authors flag all of this, use a multi model panel to blunt it, and explicitly demote HealthBench to supporting evidence while treating the blinded human review as primary. That is the correct call and it is worth saying out loud, because a lot of readers will quote the MedQA number. Frontier models write in a house style that frontier models find persuasive. Human raters are the only part of this design that escapes the loop.

Acknowledged and handled well
7
What was left out is what the clinical tools are for
What actually happened

The evaluation did not assess response latency or citation quality. The authors say so plainly in the limitations.

Why it matters

Citation quality is close to the whole value proposition of UpToDate and OpenEvidence. People do not subscribe for the prose. They subscribe because a graded recommendation traces back to a named trial, an editor's name is attached, and the whole apparatus holds up under questioning by someone who is not on your side. Grading these products on the four axes a chatbot happens to be optimized for, while omitting the axis they were built to win, is a fair test of one thing and a partial test of the product.

Title overreaches the evidence
8
One institution, twelve raters, one specialty mix
What actually happened

The queries came from a single academic health system. The graders were twelve US clinicians. The author list runs heavily toward proceduralists, with neurosurgery, cardiothoracic surgery, orthopedics, and dermatology represented.

Why it matters

What counts as a complete answer differs by specialty. In nephrology the interesting questions are longitudinal, involve four interacting problems, and turn on a judgment about how much residual function a patient will keep. A four point rubric applied to one turn of text is not built to capture whether an answer would have been right in month nine. Inter-rater concordance on model ranking was strong (Kendall's W = 0.651), so the raters agree with each other. Whether they represent all of clinical medicine is a different question.

Generalizability unresolved

The genre problem, or why this is not a paper about doctors

This study will be read as another entry in the "AI beats doctors" file. It is not one. No physician was tested against a model here. Physicians were the graders. The comparison is between two kinds of software, and the specialized software lost.

That distinction matters because the actual literature on models versus clinicians is stranger and more useful than the headlines. Two studies should be stapled to this one.

Goh and colleagues ran a randomized trial with 50 physicians on diagnostic reasoning cases. Physicians given GPT-4 scored a median 76 percent. Physicians with conventional resources scored 74 percent. The difference was not significant (P = 0.60). The model working alone scored 16 points above the conventional group. The knowledge was there. Handing it to a doctor produced almost nothing.

Bean and colleagues at Oxford went further and tested the public. Across 1,298 participants and ten clinical scenarios, the models identified the correct condition 94.9 percent of the time when tested on their own. When actual humans used those same models to work through the same scenarios, correct identification fell below 34.5 percent, no better than people using whatever they would normally use. Users did not know what to tell the model. The model answered differently depending on trivial changes in phrasing. Correct and incorrect content arrived in the same paragraph, wearing the same tone.

Every number in the Nature Medicine paper is a measurement of a model alone. The Oxford work is the measurement of a model with a person attached, and the person is where sixty points of accuracy go to die. That is not an argument against the technology. It is an argument that benchmark rank order tells you almost nothing about deployed performance, which is, to be fair, roughly the argument the authors themselves are making about vendor benchmarks. They just stopped one step short of applying it to their own.

Worth naming. The senior author discloses consulting for Google, and Gemini won all three stages. This is disclosed, it is the kind of thing that is disclosed precisely so readers can weigh it, and I see nothing in the design that suggests a thumb on the scale. But the paper's central argument is that we should stop trusting evaluations produced by people with an interest in the outcome. That standard applies in every direction.

What is genuinely useful here

I have spent eight sections on the limits, so let me be clear that I think this is a good paper and a needed one. Four things in it are worth more than the headline.

The RCQ construction is the reusable asset. One hundred real questions from real clinicians in a live deployment, de-identified under an IRB protocol, uncontaminated by anyone's training data. Any health system running an internal LLM already has this raw material sitting in its logs. Building your own version of this benchmark is the highest yield evaluation work available to a CMIO right now, because it measures your clinicians asking your questions, and because no vendor can study for it.

The error taxonomy is more actionable than the scores. Of 153 categorized failures, incomplete clinical content accounted for 55 and safety critical omission for 33. Hallucination accounted for 4. The failure mode everyone builds governance around is the rarest one in the data. Omission is the real hazard, and it is far harder to catch, because nothing on the screen looks wrong. A confident, well organized, beautifully formatted answer that leaves out the contraindication reads exactly like a correct answer. Detection strategies built to catch fabrication will not catch this.

Retrieval is not automatically an upgrade. The authors point to evidence that retrieval augmented generation can degrade performance when irrelevant material gets pulled in or poorly integrated. This is a useful correction to the assumption that grounding a model in a curated corpus is strictly additive. It is not the retrieval that helps. It is the retrieval being good, which is a much harder engineering problem than buying a corpus and a vector database.

It is a procurement fact. A subscription product at roughly $699 a year and a free ad supported product scored in the same statistical tier as the AI summary that appears above your Google results, on questions your own clinicians actually asked. That does not mean cancel the subscription. UpToDate carries editorial accountability, medicolegal defensibility, and a citation trail that none of the frontier models match, and this study did not measure any of those. It does mean the burden of proof has moved. Ask your vendor for independent evaluation on real clinical queries. If they hand you their own benchmark, you have learned something.

All of which lands at an awkward moment. In January the FDA widened enforcement discretion so that AI clinical decision support offering a single recommendation can reach the market without review, provided a clinician can independently review the basis for it. This paper argues that independent, real world evaluation of clinical AI is urgently needed. The regulatory floor moved the other way in the same six months. Whatever evaluation happens now is going to happen inside health systems, or it is not going to happen.

So What

The specialized clinical AI tools lost to general purpose models on the tests we know how to run. The tests we know how to run measure whether a doctor liked the paragraph, and the one study that put a human in the loop watched ninety five percent accuracy fall to thirty four.

Buy nothing on a benchmark. Build the hundred question set out of your own clinicians' logs, score it blinded, and count the omissions rather than the hallucinations.

Confidence: high that frontier models outperform these two products on physician rated response quality. Moderate that the ranking inside the top tier means anything. Low that any of it predicts what happens when the tool is in the room with a patient.

Sources

Primary paper: Vishwanath K, Alyakin A, Ghosh M, et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nature Medicine 2026;32:2405 to 2409. doi.org/10.1038/s41591-026-04431-5

Human in the loop failure: Bean AM, Payne R, et al. Clinical knowledge in LLMs does not translate to human interactions. arXiv:2504.18919. arxiv.org/abs/2504.18919

Oxford summary: New study warns of risks in AI chatbots giving medical advice. University of Oxford, February 2026. ox.ac.uk

Physicians plus LLM randomized trial: Goh E, Gallo R, Hom J, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open 2024. pubmed.ncbi.nlm.nih.gov/39466245

Accompanying commentary: Large Language Models, Misdiagnosing Diagnostic Excellence? JAMA Network Open 2024. jamanetwork.com

Regulatory context: FDA pulls back oversight of AI-enabled devices, wearables. STAT, January 6, 2026. statnews.com

Regulatory analysis: FDA Eases Oversight for AI-Enabled Clinical Decision Support Software and Wearables. Orrick, January 2026. orrick.com

Unregulated output: Unregulated large language models produce medical device-like output. npj Digital Medicine 2025. nature.com