The FDA Is Rethinking How Generative AI Medical Devices Should Be Tested
The Food and Drug Administration is wrestling with a deceptively simple question: how do you prove that a generative AI system is competent enough to participate in medical care?
Axios reported on August 18 that the FDA is considering a “competency-based” approach for evaluating generative-AI-enabled medical devices, drawing a rough comparison to the way physicians are evaluated on whether they can perform particular tasks safely and effectively. The idea appears in a discussion paper the agency shared with Axios and is still exactly that—a discussion. It is not a new rule, final guidance, or an FDA approval standard today.
That caveat matters, because the underlying problem is real even if the eventual answer looks completely different.
Traditional medical-device testing works reasonably well when a device behaves predictably: give it the same input under the same conditions and you expect essentially the same result. Generative AI is messier. Its output can vary, the model may be updated, and performance that looks impressive on a benchmark does not necessarily tell you how reliably it will behave across thousands of real patients and clinical situations.
The FDA has already acknowledged that problem. In an earlier request for public comment, the agency noted that many AI-enabled devices are evaluated using retrospective testing or static benchmarks, which may establish a baseline but do not necessarily predict behavior in a changing real-world clinical environment. The agency has been asking how developers and hospitals should monitor performance after deployment, including how to detect when a system begins to drift.
According to Axios, the newer discussion paper goes a step further by asking whether risk should be evaluated partly according to the job the AI is performing and the consequences if it gets that job wrong. That sounds obvious until you try to turn it into regulation. An AI system drafting routine paperwork and an AI system influencing a diagnosis can both be “medical AI,” but a bad answer carries very different stakes.
For patients, there is nothing to do because of this paper today. The more important takeaway is that regulators are beginning to move away from treating AI as merely another piece of software with a test score attached to it. In medicine, the question cannot just be whether a model is impressive. It has to be whether it is reliably good enough at the specific thing we are trusting it to do—and whether somebody notices when that stops being true.
