Buried in Appendix A of the FDA's new discussion paper on generative AI devices is a list of things a machine would have to demonstrate before it is allowed to practice. One of them is renal function based dosing. Another is recognizing clinically implausible values, including when a patient reports a number conversationally and leaves off the units. I have spent a career on the receiving end of both problems, and I can report that the humans have not solved them either.
On August 18 the Center for Devices and Radiological Health released a thirty page paper proposing that generative AI devices be evaluated the way clinicians are. Not by exhaustively testing every possible input, which for an open ended system is not a thing you can do, but by a structured assessment of knowledge and judgment followed by supervised performance in the real setting. Boards, then residency. Comments close October 19 under docket FDA-2026-N-7874.
The trade coverage framed this as the FDA weighing clinician style tests, which is right, and as a requirement, which is not. So it is worth separating what the agency actually wrote from what it is being reported to have written, and then asking the more interesting question, which is whether the analogy survives contact with the thing it is being applied to.
The competency framing is real and it is the spine of Section V. The requirement is not. The paper opens by stating it "is not intended to propose or implement policy changes," does not communicate CDRH's expectations for future marketing submissions, and, in the same breath, "is not intended to address whether the approaches discussed below are within FDA's existing legal authorities or whether new legal authorities would be necessary."
Why the distinction mattersThat last clause is not boilerplate. It is the agency declining, on page one, to answer the question that killed its last attempt at this.
Accurate. Benchmarking covers clinical knowledge, analytic capability, safety behavior, communication and generalizability. Clinical confirmation follows, and the paper lists five approaches in ascending order of rigor and patient exposure, from retrospective evaluation through shadow deployment, standardized patients, clinician adjudication of real cases, and finally a prospective study.
Accurate, and it is the most consequential sentence in the paper for anyone building one of these. Shadow deployment, in which the device runs live on real patients while its outputs are hidden from everyone and compared against what actually happened, is the approach I would expect to do most of the work.
Accurate and correctly reasoned. The unit of evaluation is "the final user-facing device, as configured and intended to be deployed for real-world use," because one general purpose model can sit under many products with different prompts and guardrails. The paper separately floats voluntary Foundation Model Master Files, under which OpenAI or Google could file a model card confidentially for sponsors to reference.
Before it gets to the exam, CDRH proposes a two axis grid for risk. One axis is what the software does, running from passive information to autonomous action. The other is what happens if you believe it and it is wrong.
To explain where a piece of software sits on the first axis, the agency wrote a four rung ladder using lisinopril.
- General information"lisinopril dosages are sometimes increased when blood pressure remains above the treatment goal"
- Connected to the user"in situations similar to this, clinicians often increase the lisinopril dosage"
- Endorsement"I recommend increasing the lisinopril dosage"
- Instruction"increase the lisinopril from 10 mg to 20 mg daily"
Every clinical informaticist who has ever argued with a vendor about whether their product is clinical decision support has just been handed a ruler. The paper then adds the part that matters more, which is that directiveness may not depend on whether the output contains the word "recommend," and that a patient facing tool does not become less directive by appending "talk to your doctor." The disclaimer does not launder the instruction. I have wanted someone with authority to say that in writing for about three years.
There is a second quiet win. The paper treats over escalation as a harm, not just under escalation. A chatbot that sends everyone with chest discomfort to the emergency department is not safe, it is expensive, and the paper says so, listing unnecessary utilization, resource burden and erosion of trust "in ways that may reduce appropriate care-seeking over time." Symmetric error accounting is rare in a safety document. Someone at CDRH has worked a triage line.
The exam itself. Ten elements across four groups. Safety (escalation, scope and boundary adherence, calibration and deferral). Clinical proficiency (knowledge, information gathering, quantitative analysis, communication). Generalizability (robustness and reproducibility, subgroup performance). Plus one element that applies only to agents, covering tool use, oversight checkpoints before irreversible actions, and resistance to prompt injection through retrieved content.
The subgroup element is unusually specific. It asks whether safety critical behavior holds across "non-standard dialects, accents, colloquialisms, and lower general literacy and health literacy levels." Not demographics as a checkbox. The actual failure mode.
Human credentialing works, to the extent it works, because the exam is the smallest part of it. The exam is wrapped in a decade of supervised practice, a license that can be revoked, a malpractice system, a hospital credentialing committee, a reputation among referring colleagues, and the specific fact that a physician who is wrong at three in the morning gets a phone call about it. The test is a filter at the front. The rest of the apparatus is what actually produces safety.
Port the exam and leave the apparatus behind and you have a score.
The benchmarks are not ready. This is the load bearing weakness, and to its credit the agency raises it as discussion question ten, asking how a sponsor should establish that benchmark performance predicts real world behavior. The empirical answer is not encouraging. A 2026 audit applying a 46 criterion framework to 56 medical LLM benchmarks found that 88 percent had no mechanism for handling data contamination, 91 percent did not evaluate a model's handling of uncertainty, and 89 percent could not test robustness to input variation. Uncertainty and robustness are two of the three things in FDA's own safety group. The tests the agency would build on do not currently measure the things the agency says it cares most about.
The comparator is a moving target. The paper suggests measuring open ended outputs against a panel of qualified clinicians reflecting the standard of care, or against "a median clinician in practice." Those are very different bars, and the second one is soft. It is also the one that will get proposed, because it is the one devices pass. Worth remembering that in a 2025 randomized trial of management reasoning, physicians using GPT-4 beat physicians using conventional resources, and GPT-4 working alone was statistically indistinguishable from the physicians using it. Median clinician performance is not a high wall. Setting it as the standard means approving devices on the grounds that they are no worse than the thing we already complain about.
Nobody loses a license. Freyer and colleagues made this point in Nature Medicine last year and CDRH cites them fairly. A clinician who fails to meet the standard faces professional, legal and reputational consequences that are personal and career ending. A device that fails faces a manufacturer with a legal department and a software update. The paper adopts the credentialing metaphor without the enforcement half, and there is no section on what happens to a device that flunks its re-benchmarking. Competency without consequence is a certificate.
The Tuesday problem. A device built on a third party foundation model can change behavior because the model developer shipped an update, with no action by the manufacturer and no notice to the hospital. CDRH sees this clearly and asks about it in question 24, floating contractual, technical and change control mechanisms. There is no good answer available. This is the genuinely novel regulatory problem in the document, and it is the one where the clinician analogy offers nothing at all. No physician has ever been silently retrained overnight by a vendor in California.
In 2017 the FDA launched the Software Precertification Pilot, which proposed evaluating the developer rather than the product. It ran for five years, involved nine companies including Apple and Verily, and was wound down in September 2022 when the agency concluded it lacked the statutory authority to run it and could not scale the model under existing law.
The new paper is a different idea. It evaluates the product, not the company, which is the right correction. But it declines to say whether the agency can do this without new authority, and that is the same wall, approached from a different angle. Congress has not moved. MDUFA VI reauthorization negotiations were underway as of last December, which is the realistic vehicle if one exists.
One footnote for the connoisseurs. Among the three papers CDRH cites as inspiration for competency based regulation, the first is by Bakul Patel, who built and ran the FDA's digital health program, Pre-Cert included, and left in 2022 to run digital health regulatory strategy at Google. He is now publishing on how the agency should regulate the category his employer sells into. Nothing improper about it. It is simply a very small field.
The competency framework is the most serious thinking the FDA has published on generative AI in medicine, and it is right about the hard part, which is that you cannot test an open ended system by enumerating its inputs. It is also borrowing a credentialing system whose exam has never been the thing that made it work.
If you build these products, comment by October 19. Question 14, on comparators, will do more to determine what clears than anything else in the document.
Sources
Primary document: FDA CDRH, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback, August 18, 2026. Full PDF and landing page. Docket FDA-2026-N-7874, comments close October 19, 2026.
Announcement: FDA press release, August 18, 2026. FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices.
The article that prompted this: PYMNTS, FDA Weighs Clinician-Style Tests for Generative AI Medical Devices, August 2026.
Cited inspirations (via PubMed): Patel B, Blumenthal D. A Novel Approach to Overseeing the Clinical Application of Generative AI. JAMA Health Forum 2026;7(3):e256947. DOI. Bergman A, Wachter RM, Emanuel EJ. A Licensure Framework for Autonomous Clinical AI. JAMA 2026;335(20):1751-1754. DOI. Freyer O, Jayabalan S, Kather JN, Gilbert S. Overcoming regulatory barriers to the implementation of AI agents in healthcare. Nat Med 2025;31(10):3239-3243. DOI.
Benchmark audit: Chen W, et al. Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models, arXiv 2508.04325v2 (revised April 2026). arXiv.
Management reasoning trial (via PubMed): Goh E, et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat Med 2025;31(4):1233-1238. DOI.
Device counts and prior program: FDA DHAC meeting summary, November 6, 2025 (Dr. Tarver, more than 1,200 AI-enabled devices authorized, none GenAI for mental health). PDF. FDA, Digital Health Software Precertification (Pre-Cert) Pilot Program, final report September 2022.