bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–08–23
twelve papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. Cochrane Evid Synth Methods. 2026 Sep;4(5): e70098
       Introduction: Artificial intelligence (AI) is a branch of technology enabling machines to emulate complex human skills; it can also entail problem-solving using bioinspired methods. It is used for automating systematic literature reviews (SLR), that is, defining a clinical question, locating relevant literature, preliminary screening, study evaluation, data extraction and analysis. Title and abstract screening is one of the most time-consuming and error-prone phases involved in developing a systematic review. While AI promises to expedite this process, adopting it faces challenges due to concerns about compatibility and transparency. This review aims to identify current evidence concerning AI use during preliminary SLR reference screening; it describes characteristics such as the different metrics used for reporting performance and how the different algorithms, pipelines, workflows or web applications are validated. AI resource users' reflections regarding SLR screening automation have also been summarized.
    Methods: A scoping review was conducted following Joanna Briggs Institute's (JBI) methodology. Its objective was to identify existing evidence regarding the use of AI resources for title and abstract screening automation. Searches were limited to articles published between 2019 and 2026. The review included primary studies reporting the development, assessment, validation, or real-world use of AI resources for screening automation, as well as systematic reviews and articles reporting experiences or recommendations for their use. Two types of data were extracted: (1) from primary studies-characteristics of AI resources and, where applicable, recommendations for their use; (2) from systematic reviews and experience-based articles, recommendations for the use of AI resources. Results included frequency descriptions, tables, figures, and a decision flowchart reflecting the number of references and articles retrieved, excluded, or included in the final analysis.
    Results: A total of 174 unique studies published between 2019 and 2026 were included in this scoping review. These were grouped into web applications (43%), model comparisons (32%), generative models (6%), pre-trained models (3%) or pipelines/workflows (15%) used for title and abstract screening in systematic literature review (SLR). Most studies came from North America. Evaluating these tools often relied on retrospective comparisons with human reviewers' work (63%), sensitivity (n = 60), and specificity (n = 62) being the most reported metric for criterion assessment and Work Saved over Sampling (n = 28) being the most reported metric for assessing their utility. Considerations concerning AI resource use focused on the need for standardized evaluation metrics, stopping criteria, study design and the data sets used, resource characteristics facilitating usability, best practice and future research areas, with the persistence of the human component in the process (n = 26) being the most pressing recommendation.
    Conclusion: The findings indicated substantial heterogeneity regarding the types of AI resources used, considerable variation concerning the metrics used for reporting performance, differences in how such metrics are defined and a clear need for standardizing reporting methods, study designs and related procedures. Although AI technologies will continue to evolve, maintaining a clear and consistent framework for interpreting research on AI resources for automating title and abstract screening can support understanding their level of maturity and facilitate informed decision-making by users.
    Keywords:  artificial Intelligence; automation; screening; systematic review methodology; validation
    DOI:  https://doi.org/10.1002/cesm.70098
  2. JMIR Form Res. 2026 Aug 20. 10 e84989
       Background: Literature reviews rely on rigorous title and abstract screening by researchers, which is time-consuming. AI-assisted literature screening tools have been proposed to improve efficiency by prioritizing titles and abstracts with the highest likelihood of meeting the inclusion criteria, thereby reducing the need to screen all records.
    Objective: This study aims to evaluate the performance of two AI-assisted screening approaches in ASReview (version 1.3; Department of Methodology and Statistics, Utrecht University) compared with manual title and abstract screening in a previously completed and published scoping review on how AI can support the quality of life in people with dementia.
    Methods: This study used a dataset of 4690 titles and abstracts from a published scoping review. The manual screening decisions of the scoping review served as the reference standard. Both ASReview approaches were applied by the same author who conducted the majority of the original manual title and abstract screening. Approach A used a simpler model with minimal prior input, whereas approach B used a more advanced model with a larger training set. Both ASReview approaches applied predefined stopping rules: (1) more than 10% of the dataset to be screened; and (2) 50 consecutive irrelevant titles and abstracts. Performance was evaluated in terms of sensitivity, specificity, precision, accuracy, and screening time. 95% CIs were calculated for sensitivity, specificity, precision, and accuracy. Agreement between manual and ASReview approaches was assessed using the Cohen κ, and differences in how manual and both ASReview approaches classified titles and abstracts were examined using the McNemar test. Performance and agreement were calculated at two levels: (1) after title and abstract screening and (2) after full-text inclusion.
    Results: Manual screening identified 283 titles and abstracts for full-text review and resulted in 30 final included studies, requiring 19 hours. Of the 4690 titles and abstracts, approach A screened 830 (17.7%) in 4.3 hours and retrieved 16 of the 30 (sensitivity 0.53, 95% CI 0.36-0.70) final included studies, whereas approach B screened 798 (17.0%) in 5.5 hours and retrieved 21 of the 30 (sensitivity 0.70, 95% CI 0.52-0.83) final included studies. Although both ASReview approaches showed high specificity and accuracy, these metrics should be interpreted cautiously because the dataset was highly imbalanced and contained relatively few relevant titles and abstracts. McNemar tests showed significant directional imbalance (P<.001): ASReview missed more manually selected titles and abstracts at level 1, whereas at level 2, ASReview more often labeled titles and abstracts not included in the final review as relevant.
    Conclusions: ASReview can support workload reduction in title and abstract screening, but the evaluated ASReview approaches did not retrieve all final included studies from the original dementia care scoping review. These findings suggest that the evaluated ASReview configurations may be insufficient for reviews in which near-complete retrieval of relevant evidence is required.
    Keywords:  AI; ASReview; dementia; information retrieval; literature screening
    DOI:  https://doi.org/10.2196/84989
  3. Ophthalmol Sci. 2026 Sep;6(9): 101296
       Purpose: To evaluate the classification performance of UveAItis, a domain-specific large language model (LLM) fine-tuned for automated title and abstract screening in systematic reviews, using retinal vasculitis as a prototype.
    Design: Comparative evaluation study embedded within a registered systematic review and meta-analysis (PROSPERO: CRD42023489232).
    Subjects: A total of 1030 randomly selected articles from an initial search of 5533 records related to retinal vasculitis.
    Methods: Articles were independently screened by 2 uveitis experts (gold standard), final-year medical students, and 3 LLMs: UveAItis (fine-tuned Generative Pre-trained Transformer [GPT]-4o), base GPT-4o, and Claude Sonnet 3.5. Screening followed a 2-question binary logic regarding human subjects and primary empirical research design. Discrepancies were resolved through expert adjudication.
    Main Outcome Measures: Classification accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and Cohen Kappa coefficient for inter-rater agreement.
    Results: UveAItis achieved the highest performance with an accuracy of 93.3%, area under the curve (AUC) of 0.887, and Kappa of 0.77. It significantly outperformed base GPT-4o (AUC: 0.805, P = 0.021), Claude Sonnet 3.5 (AUC: 0.669, P < 0.0001), and medical students (AUC: 0.585, P < 0.00001). The fine-tuned model correctly identified 65.4% of expert-included articles postconsensus, whereas students only identified 22.3%. UveAItis also demonstrated the lowest rate of ambiguous "Need Consensus" outputs (3.4%) compared to experts (16.7%).
    Conclusions: UveAItis demonstrated expert-level performance, significantly outperforming general-purpose LLMs and nonexpert human reviewers. These findings validate the potential of domain-specific fine-tuning to enhance the efficiency, scalability, and reproducibility of evidence synthesis in specialized medical fields like ophthalmology.
    Financial Disclosures: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
    Keywords:  Artificial intelligence; Large language model; Retinal vasculitis; Systematic review; Uveitis
    DOI:  https://doi.org/10.1016/j.xops.2026.101296
  4. JMIR AI. 2026 Aug 18. 5 e104210
      
    Keywords:  LLMs; PRISMA; artificial intelligence; evidence synthesis; large language models; reporting guideline; systematic literature review; transparency
    DOI:  https://doi.org/10.2196/104210
  5. Nurs Res. 2026 Aug 18.
       BACKGROUND: Large language models have accelerated the adoption of generative artificial intelligence (AI), making AI tools more widely accessible through conversational prompting. One emerging application is vibe coding, in which users use natural-language prompts to generate code and desired outputs rather than manually writing traditional code.
    OBJECTIVE: To examine AI-assisted, human-in-the-loop (HITL) vibe coding as a proof of concept for data analysis, describe its components and a proposed workflow with explicit safeguards, and present a case study illustrating its use and potential failure points.
    METHODS: We used a proposed workflow that included framing research questions, operationalizing variables, organizing project folders, documenting decisions, applying retrieval-augmented generation, and using prompt engineering techniques. We used Cursor (v1.5.11) on a limited, clean admissions data set, in which admission status was modeled as a function of the Graduate Record Exam, grade point average, and undergraduate rank. Logistic regression was generated via conversational prompts, implemented in R, and the results were compared with a published reference output on a publicly available website.
    RESULTS: AI-assisted, HITL vibe coding produced statistical codes that included schema checks, range validations, data cleaning, exploratory analyses, regression modeling, and visualization. There were mixed results of both valid and invalid outputs. Regression coefficients, p values, and model fit statistics matched the outputs posted on the published reference output website. However, an error was identified in the predicted-probability confidence interval output, which was missed during the initial review of outputs.
    DISCUSSION: While vibe coding has the potential to reduce barriers to data analysis for researchers, this case study demonstrated that it can produce both valid and invalid outputs and that foundational statistical training, knowledge, understanding, and methodological expertise remain paramount when using it. Future studies should address important empirical questions about the use of vibe coding, such as under what conditions it can be safely used in research and what kinds of errors are most commonly generated when using it. AI-assisted HITL vibe coding should be used with caution and only with structured verification and safeguards, transparent reporting, and appropriate statistical and methodological oversight.
    Keywords:  artificial intelligence; knowledge; nursing research; nursing theory; statistics as topic
    DOI:  https://doi.org/10.1097/NNR.0000000000000942
  6. JMIR AI. 2026 Aug 20. 5 e93239
       Background: Maintenance of oncology clinical practice guidelines (CPGs) is increasingly challenged by the rapid growth of trial data and therapeutic complexity. While large language models (LLMs) have shown promise in information retrieval, their utility in the rigorous, end-to-end workflow of guideline maintenance remains underexplored.
    Objective: This case study aimed to systematically evaluate the performance of frontier LLMs in supporting oncology guideline maintenance. We sought to determine their reliability in predicting necessary guideline updates based on new evidence, their accuracy in extracting data from clinical trials, and their effectiveness as automated auditors for detecting errors in established guidelines.
    Methods: Using the Onkopedia peripheral T-cell lymphoma (PTCL) guideline as a prospective case study, we tasked frontier models with deep-research modes and autonomous web-search capabilities (Gemini 2.5 Pro and GPT o4-mini-high) to predict a guideline update in August 2025 based on the 2021 version. Predictions were validated against the official 2025 revision published in October 2025. Next, we benchmarked evidence extraction accuracy across 80 pivotal trials using models of varying scale (27B-671B parameters vs frontier). Finally, we deployed a stacked LLM workflow to audit 28 recently updated Onkopedia guidelines for linguistic and content-related errors.
    Results: In the predictive task, models captured 36.7% to 40% of substantive updates, often identifying landmark approvals, but frequently overstating evidence. An independent, model-blinded rescoring yielded substantial agreement (weighted Cohen κ=0.75) and confirmed predictive accuracies of 35% to 38.3%. While frontier models demonstrated high accuracy (up to 99.2%) in extracting data from individual studies, substantially outperforming smaller open-source models, this precision declined during multisource synthesis. We observed a position-dependent performance drop in long-form generation, with GPT's endpoint accuracy dropping from 84.6% in the first half of the drafted guideline to 38.5% in the second half. As automated auditors of existing CPGs, the models successfully identified a median of 16.5 (IQR 13.8-20.3) formal errors per document and detected several clinically relevant inconsistencies (eg, invalid scoring formulas and incorrect staging definitions).
    Conclusions: LLMs currently lack the reasoning stability for autonomous guideline authoring due to deficits in complex synthesis. However, they are effective tools for high-fidelity evidence extraction and automated quality assurance, supporting a human-led, AI-augmented workflow for efficient guideline maintenance.
    Keywords:  AI; artificial intelligence; clinical practice guidelines; large language model; oncology; quality control
    DOI:  https://doi.org/10.2196/93239
  7. Eur J Neurol. 2026 Aug;33(8): e70736
       BACKGROUND: Large language models (LLMs) are entering clinical workflows for information retrieval and decision support, but concerns persist about factual accuracy and traceability to evidence. Prof. Valmed is the first CE-marked LLM-based tool for medical information retrieval in Europe, raising the question whether certification aligns with reliable performance in neurology. This study aims to assess the answer accuracy of Prof. Valmed on a neurology benchmark and compare its performance with previously evaluated commercial LLM tools.
    METHODS: Prof. Valmed (V2.0.1_2.0.0) was tested on a 130-item benchmark derived from American Academy of Neurology guidelines (65 case-based, 65 knowledge-based). Each question was prompted four times. Two raters scored responses as correct, inaccurate, wrong, or refused; disagreements were adjudicated. Modal ratings per question were used for analysis. Comparative performance against 16 LLMs or configurations previously assessed using the same protocol was evaluated.
    RESULTS: Prof. Valmed achieved 76.2% correct answers, 14.6% inaccurate, 6.9% wrong, and 2.3% refused. In pairwise comparisons, Prof. Valmed performed significantly better than 8 models, comparable to several retrieval- or whitelist-enabled tools, and significantly worse than one reasoning model with whitelisting of neurology sources. Overall, its performance was within the range of standard commercial tools.
    CONCLUSION: The CE-marked product demonstrated accuracy comparable to contemporary LLMs but did not outperform reasoning-enabled systems. CE marking ensures regulatory conformity rather than inherent performance advantage, and the certification process may limit the integration of rapidly evolving model architectures into approved systems. Future evaluations should examine reasoning quality, source use, and update feasibility across medical domains.
    Keywords:  AI; digital health; evidence‐based medicine; large language models; neurology
    DOI:  https://doi.org/10.1111/ene.70736
  8. Expert Rev Pharmacoecon Outcomes Res. 2026 Aug 17. 1-10
       INTRODUCTION: Health economic models (HEMs) provide a solid foundation for reimbursement policy decisions that shape patient access to new treatments and the allocation of scarce healthcare resources. Model development is labor-intensive and time-consuming, often requiring months of expert work. Recent advances in large language models (LLMs) prompted interest in whether artificial intelligence can support or partially automate this process, but the evidence base remains scattered and has not been mapped against the modeling workflow.
    AREAS COVERED: This review examines current applications of LLMs to health economic modeling. Five proof-of-concept studies are included and mapped to an eight-stage workflow adapted from the ISPOR-SMDM Modeling Good Research Practices framework and discussed in terms of reproducibility, validation, adaptability, and technology readiness. Published work addressed model parameterization, model implementation, reporting and quality assessment, and local adaptation, while research question design, model conceptualization, uncertainty analysis, and model validation remained unaddressed.
    EXPERT COMMENTARY: The evidence supports cautious optimism. Near-term gains are augmenting human modelers on decomposed, verifiable sub-tasks rather than pursuing autonomous end-to-end modeling, which remains distant given current reliability levels and the iterative, collaborative nature of model development.
    Keywords:  Artificial intelligence; cost-effectiveness analysis; generative AI; health economic modeling; health technology assessment; large language models
    DOI:  https://doi.org/10.1080/14737167.2026.2717289
  9. Am J Med Qual. 2026 Aug 24.
      Using artificial intelligence to identify themes in interview data from a quality improvement program evaluation produced 4 replicable themes grounded in the data. However, 2 consistently identified themes resulted from subtle misrepresentations and could have easily misled results without thorough data knowledge and output audit by the human research team.
    Keywords:  AI-assisted analysis; qualitative methods; quality improvement evaluation methods
    DOI:  https://doi.org/10.1097/JMQ.0000000000000341
  10. J Nurses Prof Dev. 2026 Jul 27.
      Nursing professional development (NPD) practitioners face increasing challenges synthesizing evidence for accredited continuing education amid expanding literature and regulatory requirements. This manuscript presents a replicable AI-assisted workflow aligned with the NPD Practice Model, NPD Scope and Standards of Practice, and ANCC accreditation criteria. Using a California implicit bias continuing education case study, it demonstrates how AI can enhance evidence synthesis and environmental scanning while maintaining NPD practitioners as the essential "human in the loop."
    DOI:  https://doi.org/10.1097/NND.0000000000001271
  11. J Vis Exp. 2026 Aug 07.
      The widespread availability of medical information on the Internet has improved access to health-related knowledge for patients; however, it has also increased exposure to inaccurate medical information. Cancer care involves complex information, making it difficult for patients to identify accurate and evidence-based sources. Large language models enable natural language interaction but frequently generate hallucinations, limiting their application in medical information delivery. Retrieval-augmented generation (RAG) has the potential to mitigate these risks by grounding responses in external reference sources. This study describes the development and evaluation of a guideline-based RAG chatbot that uses publicly available, web-based cancer clinical practice guidelines. The system extracts URLs from the table-of-contents pages of a guideline, processes page text using morphological analysis and cosine similarity based on term frequency-inverse document frequency, and constructs reference information from highly relevant pages. A large language model uses these references to generate responses constrained to guideline content. The JLCS Guidebook for Lung Cancer Patients and Families was used as the reference guideline. Question sets included both in-scope cancer types covered by the guideline and out-of-scope cancer types to evaluate response control. Medical information was successfully extracted from the table-of-contents pages, with 98% of extracted URLs containing referenceable medical content. For in-scope questions, the chatbot generated responses by summarizing guideline content, and no hallucinations were identified during manual review under the defined test conditions. For out-of-scope questions, the chatbot consistently declined to answer and indicated that the available information was insufficient. The web structure of the guideline facilitated efficient scraping and organization of reference content at the URL level. This protocol provides practical guidance for constructing artificial intelligence systems that deliver medical information using web-based clinical practice guidelines.
    DOI:  https://doi.org/10.3791/71099
  12. J Periodontal Res. 2026 Aug 22.
      Benchmarking AI chatbots using structured questions derived from the EFP S3-level guideline for stage IV periodontitis showed high clarity but variable guideline alignment and source reporting. Apparent chatbot rankings were attenuated after excluding Sources, and chatbot identity was not associated with total QAMAI score after adjustment. Rankings may therefore be influenced by retrieval and source-reporting features rather than reflecting intrinsic model capability. All outputs require clinician verification.
    Keywords:  artificial intelligence; dentistry; guidelines; language processing; large language models; periodontitis; treatment
    DOI:  https://doi.org/10.1111/jre.70166