How Should Generative AI in Healthcare Be Regulated?
AI-powered medical tools are rapidly entering clinical care. Researchers and regulators are calling for transparent, real-world testing to ensure these systems are safe, effective and accountable.
Since OpenAI released ChatGPT in November 2022, generative artificial intelligence tools have been rapidly and widely adopted in clinics. Some of the numbers are surprising.
Within four years, thousands of studies have been published describing, testing and evaluating the use of artificial intelligence in healthcare. These studies cover both general-purpose AI models and specialized systems trained on medical knowledge. By some estimates, approximately three peer-reviewed papers on AI tools used in clinical medicine are published every day.1
For many people, chatbots have become a go-to source for medical advice. Every week, more than 230 million people around the world ask health-related questions on ChatGPT.
The healthcare sector has also seen a significant increase in AI-powered medical products designed to support clinical decision-making and healthcare management. AI systems can now handle complex administrative tasks, order clinical tests, help clinicians prescribe medications,2 interpret X-rays,3 magnetic resonance imaging (MRI) scans and computed tomography (CT) images,4 and assist in diagnosing rare diseases.5
These developments raise important questions: Who regulates this innovative class of medical products? How should they be regulated? And how can the public be confident that an AI-powered health system will perform as its developers claim?
The FDA is seeking feedback on AI medical devices
The U.S. Food and Drug Administration (FDA) is seeking feedback on these questions through a discussion paper published in August. The paper considers how to regulate generative AI-enabled medical devices. Researchers are encouraged to submit their opinions by October 19. Other countries have also announced proposals for regulating AI in healthcare, including through medical-device rules.
The need for effective oversight is becoming more urgent as AI tools move beyond administrative support and begin influencing clinical decisions.
Why AI medical products require risk-based regulation
The FDA and its counterparts around the world generally approve medical products according to the level of risk they pose. In most cases, products must meet legally binding quality and safety standards, but not every product must be independently tested by the manufacturer or evaluated in real-world conditions.
For example, manufacturers of bandages may self-certify that their products meet relevant standards. Health-assessment devices such as stethoscopes often require third-party approval, although those tests can be performed in a laboratory.
By contrast, diagnostic products that influence clinical decisions must be tested in real-world conditions and, in some cases, through clinical trials similar to those used to evaluate pharmaceutical products. In general, the greater the potential harm if a product fails, the more scrutiny it requires before approval.
The FDA and other regulators are now considering whether medical products that use generative AI should undergo more comprehensive evaluation. In many cases, the answer will be yes.
AI devices that record and summarize doctor–patient conversations, for example, are more than administrative tools when their outputs are used in clinical decision-making. Some tools are also assisting doctors with disease diagnosis. These systems must be tested comprehensively and transparently before they are deployed widely.
In a comment article published in Nature Medicine in September, researchers argued that preregistered clinical trials should become the norm for AI systems used in health and medicine, just as they are for a wider range of medicines and vaccines.6
The problem with current AI evaluation benchmarks
Some AI-powered medical products are currently certified without being evaluated in real-world environments. This is concerning because these products are new, and companies have access to relatively few data sources for determining what works and what does not.
For improved stethoscope models or innovative types of bandages, decades of real-world data are available to help manufacturers benchmark new products. AI-powered devices, by contrast, may perform tasks that have never previously existed or for which clinical data are limited or nonexistent.
A systematic review published in Nature Medicine in March found that, among approximately 4,600 papers on AI tools for clinical medicine, only 23% used data from real patients. Just 19 of those studies were prospective randomized trials.1
Companies also often test AI models in one-off simulation scenarios, measuring the accuracy of a model’s decisions against the decisions made by a human doctor in the same scenario. Studies comparing specialized medical AI models with general-purpose models have produced conflicting results on whether specialized systems improve accuracy.7,8
Even an accurate AI model does not necessarily reflect the patient’s perspective. For example, companies developing these technologies do not typically ask patients whether they consent to having AI systems involved in their care.
What healthcare AI regulation can learn from self-driving cars
Some AI applications have clear effects on safety, health and well-being, yet there is less discussion about the need for real-world data. National and city governments have taken a cautious approach to self-driving cars, following principles that are not unlike those used to regulate drugs.
Self-driving vehicles have been tested for thousands of hours in laboratory conditions, simulated scenarios and on-road environments. These evaluations are necessary because errors can have serious consequences.
Real-world testing should be a requirement for clinical AI
Not every AI-powered medical device needs to be tested in a full-scale randomized controlled trial. However, the principles behind such evaluations should still apply.
Data transparency, open data and testing in real-world environments are non-negotiable. As generative AI becomes more deeply integrated into healthcare, regulators, developers and researchers must ensure that these systems are evaluated not only for technical accuracy, but also for safety, accountability and their impact on patients.
Source: www.nature.com


