Enhancing Health Policy Reviews with Machine Learning

The study by Cavallaro, Ardito, Drummond, and Ciani (2025) addresses a critical challenge in health economics and policy research: the rapidly expanding volume of literature that makes comprehensive title and abstract screening increasingly time-consuming and prone to omission. By retrospectively evaluating the open-source tool ASReview in “Simulation Mode,” the authors quantify how machine learning (ML)–assisted screening compares with traditional manual methods. Drawing on a dataset of 10 246 records (including 135 manually identified relevant studies) drawn from PubMed/MEDLINE, Scopus, and Web of Science, they simulate screening under varying levels of prior knowledge (PK) and two stopping rules—sampling and heuristic—to assess both recall and time-to-discovery metrics (Cavallaro et al., 2025).

Methodologically, the authors implement a Naïve Bayes classifier with TF-IDF feature extraction, balanced sampling, and a “Maximum” query strategy—the default configuration shown to deliver optimal performance in prior work (Ferdinands et al., 2023). They test scenarios with PK levels of 5, 10, and 15 labelled records, running 100 simulations per scenario. For the sampling criterion, they evaluate recall (percentage of relevant records found, RRF) after screening predetermined proportions of the top-ranked articles (10%, 25%, 50%, 75%), whereas for the heuristic criterion they calculate RRF based on thresholds of consecutive irrelevant records (CIRs of 100, 250, 1200, 2800) via time-to-discovery computations (Cavallaro et al., 2025).

Results indicate that ML-assisted screening achieves a median RRF of 97% after screening just 25% of the sample under the sampling rule, equating to a workload reduction of approximately 32 working days (assuming two minutes per record and an eight-hour workday). Beyond 50% screening, RRF plateaus at around 99–100% regardless of PK, suggesting diminishing returns for further manual effort (Cavallaro et al., 2025). Under the heuristic rule, performance stabilizes only at very high CIR thresholds (≥1200), but exhibits substantial variability at lower thresholds, with minimum RRF as low as 65% when CIR is 100. Increasing PK improves early-stage performance but has minimal impact once substantial screening progress has been made (Cavallaro et al., 2025).

The authors conclude that the sampling criterion—with at least 5 relevant records as PK and screening 25% of the dataset—offers a reliable and efficient strategy for title/abstract screening in health economics and policy reviews. They argue that adopting such ML tools can streamline review workflows without substantially compromising recall, although systematic reviews requiring exhaustive inclusion may still demand manual completion. Importantly, they call for regulatory bodies and HTA agencies to develop clear guidelines on ML-assisted reviews—defining approval criteria for tools, standardizing reporting of training and stopping parameters, and ensuring transparency to maintain evidence quality (Cavallaro et al., 2025).

In sum, this comparative assessment demonstrates that integrating ASReview into review protocols can achieve near-equivalent recall to manual screening while dramatically reducing effort. By framing practical recommendations—such as screening at least a quarter of records using the sampling rule and initializing with minimal prior knowledge—the study offers actionable guidance for researchers and policymakers seeking to harness ML to manage the ever-growing corpus of health policy literature (Cavallaro et al., 2025).

Reference
Cavallaro, L., Ardito, V., Drummond, M., & Ciani, O. (2025). Machine Learning–Assisted Health Economics and Policy Reviews: A Comparative Assessment. Applied Health Economics and Health Policy, 23(639–647). https://doi.org/10.1007/s40258-025-00963-y

Podcast Link: https://notebooklm.google.com/notebook/e213b090-3612-4497-a05c-cfccad90ca9c/audio

Video

Subscribe to the Health Topics Newsletter!

Google reCaptcha: Invalid site key.