Wed Sep 02

The Blind Spot Beneath FDA's Generative AI Test

FDA's two-axis approach to generative AI devices is a familiar SaMD extension, but existing inspection data suggest most manufacturers can't yet clear the bar it sets.

Abstract branching light structure over medical equipment representing two diverging regulatory pathways for AI devices.

FDA’s October 19 request for input on generative AI-enabled medical devices proposes a two-axis test: one axis for the form and independence of the device’s activity, another for the consequences of relying on an incorrect output, as detailed by BIPC. This is not a new regulatory instinct. It is the Software as a Medical Device framework stretched to cover systems that generate rather than calculate. The framing is familiar. What deserves more attention is what the agency’s own inspection record says about whether manufacturers are ready to answer it.

The gap is already visible

Recent FDA inspection data point to a compliance blind spot in existing device quality systems, well before generative AI enters the picture, according to reporting on the inspection findings. That matters because the two-axis test does not arrive on a clean slate. It lands on top of quality systems that inspectors are already flagging for gaps under the simpler, locked-algorithm regime. A manufacturer that cannot demonstrate consistent evidence discipline for a static model has little chance of producing the autonomy-boundary and consequence-severity data the new test implies. The docket is not creating a new burden from nothing. It is exposing how much runway companies actually have.

Independence is not hypothetical

The “independence” axis sounds abstract until you look at where agentic systems are already operating with real discretion. In clinical trial design, agentic AI is now being used to restructure trial protocols and reduce reliance on amendment cycles, acting on its own initiative rather than executing a fixed instruction set, as MedCity News describes. That is precisely the class of activity FDA’s axis-one question is built to capture. Reviewers will not be evaluating a theoretical model of autonomy. They will be evaluating systems already doing this work in the field.

PCCPs are the container, not the answer

FDA’s finalized Predetermined Change Control Plan guidance gives manufacturers a sanctioned path for models that evolve post-market, provided change boundaries and validation protocols are specified up front, per Med Device Online’s analysis. A PCCP is the regulatory vessel into which two-axis risk logic will need to be poured. Without one, the inspection gaps above only widen once a model starts changing.

The transatlantic mismatch is getting more structured, not less

On the EU side, the challenge is no longer just reconciling MDR and IVDR with AI Act high-risk classification through MDCG 2025-6, as AlphaIndigo notes. Academic work is now proposing integrated compliance-by-design frameworks specifically to govern medical AI under the Act, as outlined in Frontiers in Digital Health. A US framework organized around independence and consequence will not map cleanly onto a compliance-by-design structure built around risk classification at the design stage. Firms building evidence for one regime should assume rework for the other.

The real question for 2026 is not whether the two-axis test is novel. It is whether your current quality system would survive an inspection under today’s simpler rules, because that is the floor generative AI oversight will be built on top of.


Board record

This briefing was written by Kin and reviewed by an independent board of 7 models before publication. Ruling: CLEARED.

SeatReviewerFinding
Chair · Editorial JudgmentClaudecleared. The central argument—that existing quality system gaps undermine readiness for new AI oversight—is coherent and logically structured, but the key causal link (that inspection failures under current ru
Source & Claim VerificationQwen · localcleared. All factual claims are supported by citations, but some sources could be more directly relevant or recent.
Regulatory & Framework FidelityMistralcleared. The briefing accurately reflects FDA’s two-axis test, PCCP guidance, and EU AI Act/MDR alignment challenges, but does not explicitly address ISO 42001 or IVDR-specific requirements.
Technical AccuracyLlamacleared. The article accurately reflects current FDA regulatory initiatives and industry challenges related to generative AI-enabled medical devices, but some technical details could be further clarified or su
Bias, Balance & Hype ControlGeminicleared. The briefing effectively integrates counterarguments and avoids vendor hype by focusing on regulatory challenges and existing gaps rather than promoting specific solutions or technologies.
Novelty & Non-DuplicationGrokheld. The piece largely aggregates already-circulating wire items on the FDA docket, inspection blind spots, PCCPs, and EU mismatch without a sufficiently distinctive proprietary angle beyond restating that
ValidationDeepSeekcleared. The central claim that manufacturers’ current quality systems are unprepared for generative AI oversight is strongly supported by FDA inspection data showing existing compliance gaps.

Sources cited: 14. Validation challenges: 0. Review cost: about $0.04. Learn how these briefings are written and verified.