article · Proceedings of the National Academy of Sciences
Systematic literature reviews are essential for synthesising evidence in medical research, yet the manual screening of titles and abstracts requires extensive time and labour. Evaluating the capability of artificial intelligence to assist this stage, an assessment compared 18 large language models against human reviewers across three distinct systematic reviews. The models demonstrated the capacity to categorise titles and abstracts accurately into included or excluded categories, though technical limitations prevented them from classifying every single record. Across the datasets examined, using these models could decrease a human reviewer's screening workload by 33% to 93%. The findings indicate that the performance of the models is heavily influenced by how clearly inclusion criteria and review concepts are designed and formulated, highlighting the importance of precise prompting before deployment.
Conducting systematic reviews is vital for evidence-based medicine, but manually sorting through thousands of scientific papers consumes substantial time and resources. Demonstrating that artificial intelligence models can reliably perform initial screenings offers a way to accelerate health research and evidence synthesis, provided that human researchers carefully refine the review criteria used to guide the automated tools.
This research demonstrates an applied technology that could be integrated into software tools used by medical researchers, health organisations, and systematic review teams to automate initial literature screening. The models are tested and operational for workload reduction, yet real-world software adoption will require addressing technical limitations that prevent full record classification, alongside developing structured workflows to help users define precise inclusion and exclusion criteria.
AI-generated from the published abstract. Always read the original work before citing.
Systematic reviews (SR) synthesize evidence-based medical literature, but they involve labor-intensive manual article screening. Large language models (LLMs) can select relevant literature, but their quality and efficacy are still being determined compared to humans. We evaluated the overlap between title- and abstract-based selected articles of 18 different LLMs and human-selected articles for three SR. In the three SRs, 185/4,662, 122/1,741, and 45/66 articles have been selected and considered for full-text screening by two independent reviewers. Due to technical variations and the inability of the LLMs to classify all records, the LLM's considered sample sizes were smaller. However, on average, the 18 LLMs classified 4,294 (min 4,130; max 4,329), 1,539 (min 1,449; max 1,574), and 27 (min 22; max 37) of the titles and abstracts correctly as either included or excluded for the three SRs, respectively. Additional analysis revealed that the definitions of the inclusion criteria and conceptual designs significantly influenced the LLM performances. In conclusion, LLMs can reduce one reviewer´s workload between 33% and 93% during title and abstract screening. However, the exact formulation of the inclusion and exclusion criteria should be refined beforehand for ideal support of the LLMs.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1073/pnas.2411962122
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.