MARATTO

review · npj Digital Medicine

A scoping review of reporting gaps in FDA-approved AI medical devices

2024130 citationsOpen accessUniversity College Hospital, Ibadan

In plain language

An examination of 692 artificial intelligence and machine learning medical devices approved by the FDA between 1995 and 2023 reveals major reporting gaps in safety, transparency, and demographic representation. Only 3.6 percent of these approvals reported participant race or ethnicity, and over 99 percent omitted socioeconomic data entirely. Furthermore, 81.6 percent failed to document participant age. Comprehensive performance results were provided in only 46.1 percent of approvals, and fewer than two percent linked directly to peer-reviewed scientific publications detailing safety and efficacy. In addition, prospective post-market surveillance studies were present in just 9.0 percent of devices. Because these healthcare technologies are entering clinical environments without consistent demographic and evaluative reporting, there is an elevated risk that algorithmic bias and existing health disparities will be amplified.

Key takeaways

  • An assessment of 692 FDA-approved machine learning medical devices revealed widespread inconsistencies in safety and demographic reporting.
  • Only 3.6 percent of approved devices documented race or ethnicity, while over 99 percent provided no socioeconomic data.
  • Comprehensive performance study details were missing from more than half of the device approvals.
  • Fewer than ten percent of approvals contained a prospective study designed for post-market surveillance.

Why it matters

Artificial intelligence tools in healthcare risk worsening health disparities if their safety and effectiveness are not transparently verified across diverse patient groups. When regulatory records lack fundamental data on patient age, ethnicity, and socioeconomic background, clinicians and health systems cannot confirm whether a device performs equitably and safely in real-world clinical care.

Commercialisation angle

The findings inform medical device developers, technology transfer offices, and regulatory teams preparing artificial intelligence products for commercial market approval. For commercial digital health products already at or approaching market readiness, developers must address clear deficits in trial reporting by integrating thorough demographic tracking, transparent performance metrics, and prospective post-market surveillance to satisfy regulatory scrutiny and gain clinical adoption.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Machine learning and artificial intelligence (AI/ML) models in healthcare may exacerbate health biases. Regulatory oversight is critical in evaluating the safety and effectiveness of AI/ML devices in clinical settings. We conducted a scoping review on the 692 FDA-approved AI/ML-enabled medical devices approved from 1995-2023 to examine transparency, safety reporting, and sociodemographic representation. Only 3.6% of approvals reported race/ethnicity, 99.1% provided no socioeconomic data. 81.6% did not report the age of study subjects. Only 46.1% provided comprehensive detailed results of performance studies; only 1.9% included a link to a scientific publication with safety and efficacy data. Only 9.0% contained a prospective study for post-market surveillance. Despite the growing number of market-approved medical devices, our data shows that FDA reporting data remains inconsistent. Demographic and socioeconomic characteristics are underreported, exacerbating the risk of algorithmic bias and health disparity.

Research topics

  • Artificial Intelligence in Healthcare and Education
  • Biomedical and Engineering Education
  • Health Systems, Economic Evaluations, Quality of Life

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1038/s41746-024-01270-x

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.