MARATTO

article · Proceedings of the National Academy of Sciences

Transforming literature screening: The emerging role of large language models in systematic reviews

202538 citationsOpen accessZagazig University

In plain language

Systematic literature reviews are essential for synthesising evidence in medical research, yet the manual screening of titles and abstracts requires extensive time and labour. Evaluating the capability of artificial intelligence to assist this stage, an assessment compared 18 large language models against human reviewers across three distinct systematic reviews. The models demonstrated the capacity to categorise titles and abstracts accurately into included or excluded categories, though technical limitations prevented them from classifying every single record. Across the datasets examined, using these models could decrease a human reviewer's screening workload by 33% to 93%. The findings indicate that the performance of the models is heavily influenced by how clearly inclusion criteria and review concepts are designed and formulated, highlighting the importance of precise prompting before deployment.

Key takeaways

  • Across three systematic reviews, 18 large language models were evaluated on their ability to screen titles and abstracts compared to human reviewers.
  • Using large language models can reduce a single reviewer's screening workload by between 33% and 93%.
  • Model performance depends significantly on how precisely the inclusion criteria and conceptual designs are formulated.
  • Technical variations and classification failures meant the evaluated models could not process every record in the datasets.

Why it matters

Conducting systematic reviews is vital for evidence-based medicine, but manually sorting through thousands of scientific papers consumes substantial time and resources. Demonstrating that artificial intelligence models can reliably perform initial screenings offers a way to accelerate health research and evidence synthesis, provided that human researchers carefully refine the review criteria used to guide the automated tools.

Commercialisation angle

This research demonstrates an applied technology that could be integrated into software tools used by medical researchers, health organisations, and systematic review teams to automate initial literature screening. The models are tested and operational for workload reduction, yet real-world software adoption will require addressing technical limitations that prevent full record classification, alongside developing structured workflows to help users define precise inclusion and exclusion criteria.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Systematic reviews (SR) synthesize evidence-based medical literature, but they involve labor-intensive manual article screening. Large language models (LLMs) can select relevant literature, but their quality and efficacy are still being determined compared to humans. We evaluated the overlap between title- and abstract-based selected articles of 18 different LLMs and human-selected articles for three SR. In the three SRs, 185/4,662, 122/1,741, and 45/66 articles have been selected and considered for full-text screening by two independent reviewers. Due to technical variations and the inability of the LLMs to classify all records, the LLM's considered sample sizes were smaller. However, on average, the 18 LLMs classified 4,294 (min 4,130; max 4,329), 1,539 (min 1,449; max 1,574), and 27 (min 22; max 37) of the titles and abstracts correctly as either included or excluded for the three SRs, respectively. Additional analysis revealed that the definitions of the inclusion criteria and conceptual designs significantly influenced the LLM performances. In conclusion, LLMs can reduce one reviewer´s workload between 33% and 93% during title and abstract screening. However, the exact formulation of the inclusion and exclusion criteria should be refined beforehand for ideal support of the LLMs.

Research topics

  • Meta-analysis and systematic reviews
  • Artificial Intelligence in Healthcare and Education
  • Academic Writing and Publishing

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1073/pnas.2411962122

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.