Campbell Syst Rev. 2026 Sep;22(3):
18911803261469813
Background: Natural language processing (NLP) techniques offer promising solutions for semi-automating the time-consuming process of abstract screening in systematic reviews. The exponential growth of published literature has created significant bottlenecks, with review teams manually assessing thousands of abstracts over weeks to months. Single reviewers can miss 5-13% of relevant studies, necessitating dual screening that further increases workload. Advances in artificial intelligence, including deep learning models such as BERT and its successors, show potential for automating this critical step, but comprehensive evidence on optimal approaches, performance, and practical feasibility remains limited.
Objectives: This systematic review aimed to assess techniques, performance, and feasibility of NLP approaches for title and abstract screening by characterizing the range of NLP methods used, summarizing performance on key metrics like workload reduction and recall, evaluating real-world implementation feasibility, and identifying research gaps and future directions.
Search Methods: We searched PubMed, Web of Science, Embase, CINAHL, The Cochrane Library, Scopus, and gray literature sources from inception to December 2024. The search strategy, developed with an information specialist and peer-reviewed using PRESS guidelines, targeted keywords related to natural language processing, machine learning, abstract screening, and systematic reviews. Additional sources included conference proceedings, preprint servers, reference lists, forward citation tracking, and expert consultation.
Selection Criteria: We included primary studies of any design describing development or evaluation of NLP techniques for automating title and abstract screening in evidence syntheses. Eligible studies reported on NLP methods, screening performance (workload reduction, recall, precision), or implementation feasibility. Studies using only rule-based approaches without machine learning, systematic reviews of NLP methods, commentaries, and conference abstracts were excluded. No language or date restrictions were applied.
Data Collection and Analysis: Two reviewers independently screened titles, abstracts, and full texts using Covidence software, with disagreements resolved through discussion. Data extraction covered study characteristics, NLP techniques, training approaches, performance metrics, and feasibility considerations. Risk of bias was assessed using a modified ROBIS tool. Given diverse techniques and outcomes, we conducted narrative synthesis following SWiM guidelines, grouping studies by NLP approach.
Main Results: From 4,105 records, 19 studies met inclusion criteria, with 68.4% published since 2023, reflecting rapid field advancement. Studies employed diverse approaches from traditional machine learning (Support Vector Machines, Random Forests) to advanced deep learning models, particularly BERT variants. Most achieved >90% recall with workload reductions of 13-96%, representing substantial time savings. Deep learning models with transfer learning consistently outperformed traditional approaches. However, implementation faced significant barriers including requirements for high-quality training data, specialized computational resources, technical expertise, and user-friendly interfaces. Performance was generally better for targeted reviews with lower inclusion prevalence.
Authors’ Conclusions: NLP techniques, especially deep learning with transfer learning, show substantial promise for semi-automating abstract screening with potential for large workload savings while maintaining high recall. However, challenges remain regarding training data quality, computational requirements, technical expertise needs, and user-centered design. Realizing full potential requires interdisciplinary collaboration to develop reliable, generalizable tools integrating seamlessly with human expertise and existing workflows. Future priorities include creating standardized datasets, conducting prospective evaluations, developing user-friendly interfaces, and establishing implementation best practices to revolutionize evidence synthesis efficiency.
Keywords: abstract screening; deep learning; evidence synthesis; machine learning; natural language processing; systematic review automation