The US Food and Drug Administration is accepting comments until 19 October 2026 on a discussion paper about generative-AI-enabled medical devices. The paper outlines a possible two-axis risk framework, a premarket competency assessment combining non-clinical benchmarks with clinical confirmation, and risk-proportionate post-market monitoring. It also raises questions about foundation models and agentic systems. The FDA announcement describes the proposal and docket.

Why a static accuracy score is insufficient

A conventional model may produce a bounded classification from a defined input. A generative system can create free-form explanations, summarize records, ask questions, recommend next steps or use tools. Its output can change with model updates, prompts, retrieved information and workflow context. A benchmark may show capability under test conditions without proving safe behavior across real users and edge cases.

The FDA's competency concept is inspired at a high level by professional training: test knowledge and performance before clinical use, then confirm that the device performs as intended in the actual care context. This remains a discussion proposal, not final guidance or an approval pathway that guarantees market access.

Risk depends on use, autonomy and reversibility

A tool that rewrites an appointment reminder is not equivalent to an agent that prioritizes patients or recommends clinical action. Assessment should consider the severity of a wrong output, how directly it changes care, whether a qualified person reviews it, the diversity of affected patients and how easily the action can be reversed.

Agentic systems add another layer because they can call tools, retrieve data and execute steps. Safe design must evaluate permissions and workflow behavior, not only the language model. A correct recommendation can still cause harm when sent to the wrong system or executed without authorization.

A readiness pack for healthcare organizations

Document the intended use, prohibited use and human owner for every AI function. Build a representative test set covering languages, clinical contexts, vulnerable groups and known failure modes. Record model, prompt, retrieval source and tool versions so a change can be traced.

Before launch, define escalation thresholds, override controls and a rollback plan. After launch, monitor harmful-error rate, unsupported statements, correct escalation, subgroup performance, overrides, complaints and time to investigation. Revalidate after meaningful model or workflow changes rather than assuming the original evidence still applies.

Commercial and GCC implications

Saudi and UAE health providers evaluating imported AI products should ask vendors for intended-use boundaries, local-language evidence, update policy, data location, audit logs and incident response. US regulatory status can inform diligence but does not replace local requirements from health, data and medical-device authorities.

Marketing teams must not convert an experimental capability into a clinical promise. Claims should match the evaluated use, identify the role of clinician oversight and explain that an AI output may support rather than replace professional judgment.

Karim's strategic takeaway

Treat governance as part of product value. A healthcare AI service that can show its limits, evidence, approvals, logs and monitoring will be easier to trust and sell than a more impressive demo with unclear responsibility. Start with a low-risk workflow, prove control and expand only when the evidence supports it.