AI-Assisted Systematic Literature Reviews: A GPT-4 System

This article, titled “Enhancing systematic literature reviews with generative artificial intelligence: development, applications, and performance evaluation“, presents a novel approach to streamlining the traditionally arduous process of Systematic Literature Reviews (SLRs). Authored by Ying Li and a team of researchers from Regeneron Pharmaceuticals, Inc., and IMO Health, Inc., the paper focuses on the development and validation of a large language model (LLM)-assisted system designed specifically for conducting SLRs, particularly for Health Technology Assessment (HTA) submissions.

The core motivation behind this research stems from the inherent challenges of traditional SLRs. These reviews, which represent the highest form of evidence in evidence-based medicine and are crucial for informing clinical practice and healthcare policy, are notoriously time-consuming, labor-intensive, and costly. They can take anywhere from 6 to 16 months, or even up to two years for the most comprehensive ones. A significant bottleneck in this process is literature screening, which is prone to reviewer disputes, human errors like fatigue and bias, and is difficult to replicate or adjust due to the substantial human effort required to review thousands of citations. Furthermore, the growing concept of “living SLRs,” which require continuous updating of literature, makes traditional methods resource-intensive.

To address these critical issues, the authors propose an AI-assisted SLR (AI-SLR) system utilizing GPT-4 in a zero-shot setting. This innovative system is structured into five distinct modules:

  • Literature search query setup: For acquiring abstracts from databases like PubMed and/or Embase.
  • Study protocol setup: Where inclusion/exclusion criteria are defined using the Population, Intervention/Comparison, Outcome, and Study Type (PICOs) framework.
  • LLM-assisted abstract screening: GPT-4 performs the initial abstract screening based on the PICOs criteria, with human users reviewing a sample and providing feedback for iterative refinement of criteria.
  • LLM-assisted data extraction: GPT-4 extracts relevant information such as study details, patient characteristics, and predefined outcomes from the included abstracts.
  • Data summarization: Consolidates screening decisions, exclusion reasons, and extracted data into structured reports.

A key strength of this system is its human-in-the-loop design. This feature allows for real-time adjustments and feedback from subject matter experts, providing them with greater control through prompt adjustment and iterative refinement of the PICOs criteria based on performance metrics. Unlike many previous automation attempts that relied on labor-intensive labeling and training data, this system leverages GPT-4’s in-context learning capabilities, thus eliminating the need for manually annotated training data, making it both time-efficient and highly generalizable across various research topics. The authors highlight its potential to streamline SLRs, significantly reducing time, cost, and human errors, while enhancing evidence generation for HTA submissions. Resource savings of 61-79% for the title and abstract screening phases have been identified when using machine learning methods.

The system’s performance was rigorously evaluated using four datasets, including studies on relapsed and refractory multiple myeloma (RRMM) and advanced melanoma. The results demonstrated relatively high performance across all evaluation sets. For abstract screening, the system achieved an average sensitivity of 90%, an F1 score of 82, an accuracy of 89%, and a Cohen’s κ of 0.71, indicating substantial agreement with human reviewers. Its sensitivity (ranging from 82-97%) is particularly noteworthy, aligning with clinical SLR practices that emphasize inclusiveness at the abstract screening stage. In identifying specific exclusion rationales, the system attained high accuracies of 97% (for RRMM) and 84% (for advanced melanoma), and F1 scores of 98 and 89 respectively. For data extraction, the system achieved an F1 score of 93. It performed particularly well in extracting study details and patient characteristics (100% F1 score for RRMM).

The authors acknowledge areas for future work, including assessing the system’s performance in medical fields beyond oncology, evaluating the effectiveness of the human-in-the-loop mechanism, and extending the system to handle screening and data extraction from full-text articles. They also plan to incorporate reinforcement learning for automated refinement of PICOs criteria and benchmark time savings in real-world settings.

In conclusion, this article presents a generalizable, end-to-end LLM-based AI-SLR system that is a significant step towards modernizing systematic literature reviews. It is notable for being, to the authors’ knowledge, the first to use PICOs criteria to instruct an LLM, further enhancing its applicability in evidence-based medicine.

Reference: Li, Y., Datta, S., Rastegar-Mojarad, M., Lee, K., Paek, H., Glasgow, J., Liston, C., He, L., Wang, X., & Xu, Y. (2025). Enhancing systematic literature reviews with generative artificial intelligence: development, applications, and performance evaluation. Journal of the American Medical Informatics Association, 32, 616–625. https://doi.org/10.1093/jamia/ocaf030

Video

Podcast Link

https://notebooklm.google.com/notebook/97b97654-ed0a-400d-b9ce-6fd4d27c4b83/audio

Subscribe to the Health Topics Newsletter!

Google reCaptcha: Invalid site key.