Blog - Kapstone Medical

The Medical Device AI Regulatory Problem Nobody Has Solved

Written by Kara Johnson | Sep 16, 2026, 4:15:06 PM

How generative AI challenges medical device regulation — validation, risk ownership, and the future of Regulatory Affairs

Medical device regulation has largely been built around a single premise: when a device is used for its intended purpose, the manufacturer can define, characterize, and validate how that device is expected to behave. That premise becomes much harder to apply when a medical device incorporates generative AI, whose outputs can be probabilistic, variable, and capable of evolving over time.

The central question is no longer just “how do we validate the software?” but “how do we establish a reasonable assurance of safety and effectiveness when the behavior of the technology is not completely deterministic?”

— Kara Johnson, VP of Regulatory

This article walks through four dimensions of that problem — how existing frameworks strain under generative AI, what validation means when outputs vary, who owns AI-related risk, and how AI may reshape the Regulatory Affairs profession itself.

Key Takeaways:

  • Traditional device regulation assumes reproducible behavior; generative AI breaks that assumption. When the same input can produce different outputs, the regulatory question shifts from “does the software produce the specified output?” to “does the system stay within an acceptable range of safe and effective performance?”
  • Validation may become a lifecycle activity, not a one-time pre-market event. For AI-enabled devices that can change or learn, post-market performance monitoring can become integral to validation — and the FDA’s finalized Predetermined Change Control Plan (PCCP) framework offers a way to pre-authorize bounded modifications without a new marketing submission.
  • AI risk is inherently cross-functional. A single AI performance issue can simultaneously be a data-quality, bias, clinical, verification, risk-management, and labeling issue — so accountability cannot sit with Regulatory, Quality, Engineering, or Clinical alone.
  • Third-party AI does not transfer responsibility. When a model is built or maintained by a supplier, the manufacturer still remains responsible for demonstrating that the finished device meets applicable regulatory and quality requirements, which raises the bar on supplier controls.
  • AI will likely elevate Regulatory Affairs rather than eliminate it. It shifts professionals from information processing — searching, compiling, first drafts — toward higher-value regulatory judgment, strategy, and defensible decision-making.

 

Part 1 — Why traditional medical device regulation struggles with generative AI

Traditional software validation generally addresses a well-bounded set of questions: What are the intended inputs? What outputs should the software generate? Can we verify that the software consistently performs as specified? What happens when the software behaves unexpectedly?

Those questions get much harder to answer when the same prompt, clinical scenario, or patient data can produce different outputs — and when the underlying model, data, or performance characteristics may change over time. A static risk-management framework was not designed for a moving target.

This creates a regulatory challenge that goes well beyond “validating the software.” The fundamental question becomes: how do you establish a reasonable assurance of safety and effectiveness when the behavior of the technology is not completely deterministic?

That single question has implications across virtually the entire medical device regulatory framework, including:

  • Design controls
  • Risk managementVerification and validation
  • Clinical evaluation
  • Labeling
  • Change control
  • Post-market surveillance

It is worth being precise about what the challenge actually is. It is not simply “how do we regulate AI.” It is whether regulatory frameworks developed largely around defined and reproducible device behavior can adequately address technologies whose outputs may be probabilistic, variable, and potentially evolving. The FDA has not thrown out its existing framework — AI-enabled device software functions are still regulated through the established pathways (510(k), De Novo, and PMA where applicable) — but the agency, along with international partners, are actively developing new mechanisms to fit these technologies into that framework. We are only beginning to understand what that means in practice.

Part 2 — How do you validate a device that doesn’t always give the same answer?

Here is a regulatory question I find particularly interesting: what does “validation” mean when the output isn’t always predictable?

For traditional medical device software, we establish defined requirements and expected behavior, then verify and validate the software against those requirements under specified conditions. But imagine a device using generative AI to assist a physician. The physician enters the same clinical information twice, and the system produces two different recommendations — neither of which is necessarily wrong. So which output represents the “correct” result?

This creates a fundamentally different validation challenge. Instead of asking “does the software produce the specified output?”, we may need to ask “does the system consistently remain within an acceptable range of safe and effective performance?” That shift moves the regulatory conversation toward concepts such as:

  • Performance boundaries

  • Acceptable variability

  • Failure modes and edge cases

  • Human oversight

  • Dataset quality and representativeness

  • Bias

  • Model drift

  • Ongoing performance monitoring

That last point may be the most consequential. For conventional software, a manufacturer can establish a validated state and then evaluate subsequent changes against that baseline. But what happens when the technology itself can change, learn, or evolve based on new data or conditions? The traditional relationship between validation, change control, and post-market surveillance becomes far more complicated.

This raises a genuinely hard question: could post-market surveillance become an integral component of the validation strategy for certain AI-enabled medical devices? Not because manufacturers should be able to “validate in the field,” but because real-world performance may reveal information that cannot be fully characterized before commercialization. If so, validation starts to look less like a single pre-market event and more like an ongoing lifecycle activity.

This is precisely the problem the FDA’s Predetermined Change Control Plan (PCCP) framework is designed to address. Finalized in the guidance Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions, a PCCP lets a manufacturer describe, in the original marketing submission, the specific modifications an AI-enabled device is expected to undergo, the methodology to develop and validate those modifications, and an assessment of their impact. When the FDA authorizes a PCCP, the manufacturer can implement those pre-specified changes without a new marketing submission — provided the changes stay within the bounds of the authorized plan. The framework draws on statutory authority added by the Food and Drug Omnibus Reform Act (FDORA) of 2022, and it is complemented by the FDA/Health Canada/MHRA Good Machine Learning Practice (GMLP) guiding principles, which set expectations around representative data, independent test datasets, and post-deployment monitoring.

Even with a PCCP, a central question remains: where should the line fall between acceptable ongoing learning and a change significant enough to require additional regulatory review? I suspect that will become one of the most challenging questions in AI-enabled medical device regulation.

Part 3 — Who owns AI risk?

As medical devices increasingly incorporate AI, an important question emerges: who is responsible for identifying, evaluating, and controlling AI-related risk? The easy answer is “the manufacturer.” The more useful answer is that this responsibility cannot sit with any single function.

Identifying and managing AI-related risk is unlikely to be the job of Regulatory Affairs, Quality, Engineering, or Clinical alone. It requires a coordinated, cross-functional approach, because each discipline sees only part of the picture:

  • Software / Engineering understands the software architecture, model characteristics, performance limitations, and technical controls.

  • Clinical evaluates whether the device’s outputs are clinically appropriate and what impact an incorrect output could have on a patient or user.

  • Regulatory Affairs evaluates applicable requirements, regulatory strategy, and the evidence necessary to demonstrate safety and effectiveness.

  • Quality establishes the processes and controls that ensure the device is developed, manufactured, changed, and monitored in accordance with quality system requirements.

  • Risk Management provides the framework for identifying hazards, evaluating risks, implementing controls, and determining whether residual risk is acceptable.

The difficulty is that AI-related risks routinely cross all of these disciplines at once. Consider an AI-enabled device that performs well overall but performs less effectively for a particular patient population. How should that be characterized? It could be a software performance issue, a data-quality issue, a bias issue, a clinical performance issue, a verification and validation issue, a risk-management issue, or a labeling issue. The honest answer is that it may be more than one — and the appropriate assessment depends on the intended use, device characteristics, identified hazards, and potential impact on safety and effectiveness.

The situation becomes more complex when the AI is developed or maintained by a third-party supplier. The manufacturer may not control the underlying model development, training data, or architecture — yet the manufacturer remains responsible for ensuring that the finished device meets applicable regulatory and quality requirements. Under the FDA’s Quality Management System Regulation (QMSR), which incorporates ISO 13485 by reference, that responsibility runs through supplier controls, design controls, and change management. AI cannot be treated as just another purchased software component.

This is why AI-enabled devices will require stronger cross-functional governance and clearly defined responsibilities across the product lifecycle. Introducing AI can affect design and development, risk management, verification and validation, supplier controls, change management, clinical evaluation, labeling, and post-market surveillance. Ultimately, the regulatory question is not simply who “owns” AI risk — it is whether the manufacturer can demonstrate that AI-related risks have been identified, adequately evaluated, appropriately controlled, and monitored throughout the device lifecycle. That is where Regulatory, Quality, Engineering, Clinical, and Risk Management have to work together.

Part 4 — Could AI ultimately change the regulatory profession itself?

The first three parts examine how AI challenges the way we regulate medical devices. There is a further question worth asking: how will AI change the practice of Regulatory Affairs itself?

AI systems are increasingly capable of tasks that occupy a great deal of regulatory time — identifying potential predicates and comparable devices, analyzing large volumes of FDA decisions, comparing requirements across markets, reviewing risk-management documentation, spotting gaps in submission content, monitoring standards changes, drafting portions of submissions, analyzing complaint and post-market data, and synthesizing literature. Many of these are already practical tools for regulatory professionals.

But there is a critical distinction between processing regulatory information and making a regulatory decision:

  • AI can identify potentially relevant predicates. It cannot assume responsibility for determining whether a predicate provides an appropriate basis for substantial equivalence.

  • AI can flag inconsistencies in a risk analysis. It cannot independently determine whether the overall risk-management strategy is appropriate for the device and its intended use.

  • AI can identify applicable regulatory requirements. It cannot replace the professional judgment needed to interpret and apply those requirements to a specific product and strategy.

And ultimately, someone remains accountable for the regulatory position the manufacturer takes.

That accountability may fundamentally reshape the role. Less time may be spent searching for information, compiling data, formatting documents, performing repetitive reviews, and generating first drafts. More time may be spent developing regulatory strategy, assessing the quality and relevance of evidence, interpreting requirements, evaluating uncertainty, challenging assumptions, managing regulatory risk, communicating implications to leadership, and making defensible decisions.

In that sense, the most significant impact of AI is probably not the elimination of Regulatory Affairs. It is the evolution and refinement of the profession from information processing toward higher-value regulatory judgment and strategic decision-making — which could be a very positive development. It also raises a real question: if AI increasingly performs the work that has traditionally filled much of an RA professional’s day, will our definition of what it means to be a “regulatory professional” need to change?

The bottom line

The challenge of regulating AI-enabled medical devices is not that the technology is unregulatable. It is that frameworks built around defined, reproducible behavior must adapt to technologies whose outputs may be variable and evolving. That adaptation is already underway — through mechanisms like Predetermined Change Control Plans, Good Machine Learning Practice, and a total-product-lifecycle mindset — but the hardest questions about validation, risk ownership, and accountability are still being worked out. Manufacturers that build strong cross-functional governance now, and treat AI as a lifecycle responsibility rather than a software feature, will be best positioned as the regulatory expectations continue to take shape.

At Kapstone Medical, our regulatory team helps device companies — including software-, algorithm-, and AI-enabled developers — build regulatory and quality strategies that hold up under exactly these questions. As an ISO 13485-certified, FDA-registered, QMSR-compliant single-source partner with more than 100 successful regulatory submissions, we work alongside engineering, clinical, and quality teams to define pathways, structure evidence, and manage AI-related risk across the product lifecycle.

Frequently Asked Questions

How does the FDA regulate AI-enabled medical devices?

The FDA regulates AI-enabled medical devices through its existing device framework for software functions, using the established pathways — 510(k), De Novo, and PMA — where applicable. It supplements that framework with tools tailored to AI, most notably the Predetermined Change Control Plan (PCCP) for pre-authorizing certain model modifications, and the Good Machine Learning Practice (GMLP) guiding principles. AI-enabled device software functions are evaluated using a total-product-lifecycle approach rather than a one-time review.

What is a Predetermined Change Control Plan (PCCP)?

A PCCP is FDA documentation, included in a marketing submission, that specifies the AI/ML modifications a device is expected to undergo, the methodology to develop and validate those modifications, and an assessment of their impact. Once the FDA authorizes a PCCP, the manufacturer can implement the pre-specified changes without submitting a new marketing application, as long as the changes stay within the authorized plan. The FDA finalized this framework in December 2024 in a dedicated guidance, Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions.

Why is validating a generative AI medical device so difficult?

Traditional validation confirms that software produces a specified output under specified conditions. Generative AI can produce different outputs from the same input, and its performance can drift or evolve over time, so there may be no single “correct” output to validate against. As a result, validation shifts toward demonstrating that the system remains within an acceptable range of safe and effective performance, and it often extends into ongoing post-market monitoring rather than ending at launch.

What is Software as a Medical Device (SaMD)?

Software as a Medical Device (SaMD) is software intended to be used for one or more medical purposes that performs those purposes without being part of a hardware medical device. Many AI- and algorithm-driven diagnostic, monitoring, and clinical-decision-support products fall into this category, and they are subject to medical device regulatory requirements based on their intended use and risk.

What is Good Machine Learning Practice (GMLP)?

GMLP is a set of ten guiding principles issued in 2021 by the FDA, Health Canada, and the UK’s MHRA to promote safe, effective, and high-quality machine-learning-enabled medical devices. The principles are not legally binding, but the FDA expects AI/ML submissions to address them — particularly representative training data, independent test datasets, and post-deployment performance monitoring.

Does ISO 13485 or the FDA’s QMSR cover AI-enabled devices?

Yes. AI-enabled devices are subject to the same quality system expectations as other medical devices, and in the US the Quality Management System Regulation (QMSR) — which incorporates ISO 13485 by reference and takes effect February 2, 2026 — applies. For AI, the most relevant areas typically include design controls, risk management, supplier controls, and change management, since these govern how the model is developed, validated, monitored, and modified.

Who is responsible for AI-related risk in a medical device?

The manufacturer is ultimately responsible, but managing AI risk requires a cross-functional effort across Regulatory Affairs, Quality, Engineering, Clinical, and Risk Management. A single AI performance issue can simultaneously be a data-quality, bias, clinical, validation, risk-management, and labeling issue, so no single function can own it alone. The regulatory test is whether the manufacturer can demonstrate that AI-related risks were identified, evaluated, controlled, and monitored across the device lifecycle.

If our AI model is built by a third-party supplier, are we still responsible?

Yes. Even when a third party develops or maintains the model, training data, or architecture, the medical device manufacturer remains responsible for ensuring the finished device meets applicable regulatory and quality requirements. This makes supplier controls, clear contractual responsibilities, and transparency into model development and performance essential parts of the risk strategy.

Do we need a new FDA submission every time our AI model changes?

Not necessarily. Changes that fall within an FDA-authorized PCCP can be implemented without a new marketing submission, provided they match the pre-specified modifications and protocols. Changes that fall outside the authorized plan — for example, a new intended use or a modification not contemplated by the PCCP — generally require additional regulatory review.

We’re an AI or digital health company based outside the US — how do we enter the US market?

Non-US manufacturers of AI-enabled devices follow the same FDA pathways as other device makers (typically 510(k), De Novo, or PMA), must appoint a US Agent and complete establishment registration, and should plan their quality system against the QMSR. A CE Mark or other foreign approval does not substitute for FDA review, though much of the underlying evidence can often be repurposed. Kapstone Medical regularly helps international companies map this route; see our companion guide, CE Mark in Hand, Now What?.

Additional Resources