Explore our Topics:

FDA floats “board exam” model for regulating GenAI medical devices, asks for industry comment

A new discussion paper proposes testing GenAI devices like med students, comparing them to average clinicians, and monitoring for drift after launch.
By admin
Aug 25, 2026, 9:53 AM

In 2023, headlines were abuzz with the startling news that an early version of ChatGPT had achieved a passing score on the US Medical Licensing Exam, heralding the imminent obsolescence of human clinicians.

Three years later, human care providers are in no danger of losing their jobs. But more and more of them are being joined by AI colleagues that can augment and optimize the important work that they do – sometimes to an eerily advanced degree.

This growing category of AI clinical assistance tools isn’t just changing the provider workflow. It’s also challenging regulators to rethink traditional methods of assessing software tools for safety, bias, and effectiveness.

Now, the FDA is considering whether established risk-based frameworks for software evaluation remain sufficient in the era of generative AI solutions – or whether these tools need a more nuanced approach.

In a newly released discussion paper, the FDA highlights the central problem facing regulators: When GenAI can produce a basically infinite number of outcomes based on an uncountable number of variables and calculations, how is it possible to prove that enough of the outcomes will be safe and effective enough of the time?

The answer may lie in treating the software more like a person: establishing an acceptable baseline of competency and continuing to monitor outputs over time to make sure they stay aligned with its defined clinical role.

The need for a new approach to software evaluation

In the paper, the FDA explicitly states that “evaluation approaches developed for software with bounded inputs and fixed outputs may not be appropriate for GenAI-enabled devices,” because unlike non-GenAI tools that can generally be held up against a representative set of inputs that produce an expected set of outputs, GenAI simply doesn’t function the same way.

Large language models (LLMs) enable these tools to be highly variable and adaptive, with the same prompt potentially producing different results based on a number of explicable and non-explicable factors.

Therefore, the agency says, “it may be unreasonable to evaluate every conceivable input that the device might encounter during deployment,” the same way it’s not possible to test a medical student on every single type of clinical situation she may encounter throughout her entire career before handing her a diploma.

Instead, “it will be important for manufacturers to consider how to effectively evaluate and monitor the GenAI-enabled device, and the underlying foundation model as applicable, for its specific intended use in a way that can ensure that device accuracy, relevance, and reliability are maintained, once deployed,” the paper says.

Doing so will require a different type of AI regulation: one that could possibly incorporate exam-style testing with supervised piloting, periodic reevaluation, and a defined scope of practice.

What could a clinician-style evaluation framework look like?

The paper outlines a possible two-part premarket evaluation system for GenAI tools, starting with the equivalent of a “board exam” that provides a benchmark for safety and accuracy.

The exam would cover five dimensions of competence, including clinical knowledge, analytic capabilities, safety behaviors, communication, and generalizability. For example, the tool would need to sufficiently prove that it can recognize safety issues and appropriately escalate those issues to human users, perform as expected across different populations, and reach accurate conclusions using the most applicable data points available to it.

Upon passing initial tests, potentially conducted by independent adjudicators, the GenAI tools would need to be exposed to some degree of real-world testing. The specific design of the testing may be based on risk, aligning well with the FDA’s existing risk-based approach to medical device evaluation, and could range from limited retrospective analysis to a full shadow deployment in a real-world clinical environment.

Ongoing monitoring may be appropriate after initial market authorization, too, the FDA suggests. GenAI models can experience “drift” due to changing patient populations, data environments, or underlying model components, which may require long-term competency surveillance and periodic re-benchmarking.

What does “good enough” mean for GenAI?

The paper presents an interest conundrum around setting guardrails for accuracy. Human clinicians don’t face a strict “three strikes and you’re out” limit when it comes to differential diagnosis, but should GenAI tools have similar leeway to iterate and adapt, or must it present the right answer the first time, every time?

The middle ground may be the most effective way forward.

“For many GenAI-enabled devices, performance might be compared to that of a panel of qualified clinicians whose consensus reflects the applicable standard of care, or to that of a median clinician in practice,” the paper explains.

“With such an approach, the relevant question might be whether the device performs as well as, or better than, a reference clinician on the same task, while recognizing that additional questions may arise regarding whether generalist or specialist physicians should be used for comparison.”

The right comparators for a given use case might depend on how the tool is actually used, including “the combined performance of the clinician and device working together as a human-AI team versus the device working in a fully autonomous workflow, depending on the intended use,” the FDA continues.

However, it’s unclear how this “average clinician” approach would dovetail with emerging questions about legal accountability and liability for harm, or how it might satisfy patient uncertainty around the safety and accuracy of AI tools being used in their care.

A request for feedback on important regulatory questions

The paper also serves as a request for industry feedback on these crucial questions about the future of GenAI regulation. Specifically, the FDA is interested in hearing from stakeholders on key issues including:

Is competency-based assessment actually the right framework for evaluating GenAI?

Are there opportunities for the FDA accept greater pre-market uncertainty in exchange for more robust post-market monitoring?

What role should healthcare institutions, clinicians, developers/manufacturers, and independent third-party evaluators play in the evaluation processes? What about AI agents?

What happens if the scope of future modifications for GenAI tools can’t be fully pre-specified during testing?

What responsibility do health systems have after deployment of a GenAI tool? What about vendors who don’t own the foundation models upon which their tools are built? How can the FDA establish accountability without restricting innovation?

Interested parties can submit their feedback on these and other questions through October 19, 2026.


Show Your Support

Subscribe

Newsletter Logo

Subscribe to our topic-centric newsletters to get the latest insights delivered to your inbox weekly.

Enter your information below

By submitting this form, you are agreeing to DHI’s Privacy Policy and Terms of Use.