bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–10–04
eleven papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. J Clin Epidemiol. 2026 Sep 28. pii: S0895-4356(26)00417-8. [Epub ahead of print] 112541
       OBJECTIVE: Manual screening of titles, abstracts, and full texts for large-volume knowledge synthesis requires significant time investment. Machine-assisted tools offer solutions for facilitating screening. Most applications to date have focused on actively reordering citations or suggesting screening decisions for reviewers to prioritize relevant studies, or on automating exclusions based on narrowly defined eligibility criteria. However, few tools extend to full-text screening, and among those that report, a majority have largely been developed with little engagement of knowledge users. Knowledge user input has the potential to foster the relevance of review outcomes and uncover nuances such as broader community participation in research, that is often variably defined and applied. We aimed to describe a pragmatic machine-assisted screening workflow that extends to full-text screening and engages knowledge users throughout.
    STUDY DESIGN AND SETTING: We implemented a pragmatic machine-assisted workflow for screening in a large-volume scoping review of community participatory approaches in infectious disease mathematical modeling. The review considered studies on any specified population, modeling any infectious disease, and where community participation was reported in at least one modeling step. The workflow was co-developed with community researchers, representing community-based organizations, who contributed to defining eligibility criteria and screening. First, we screened titles and abstracts using a supervised random forest model. We then triaged full texts based on keywords and made final classifications (inclusion/exclusion) using ridge-penalized logistic regression. We evaluated the machine learning model performances using held-out test sets and bootstrapped 95% confidence intervals.
    RESULTS: Of 16,629 studies identified in the search, 9,841 unique citations underwent title and abstract screening and 2786 citations underwent full-text screening. Model-based title and abstract screening achieved sensitivity of 86.7% (95% CI: 84.1-89.4) and area under the curve (AUC) of 86.4% (95% CI: 84.4-88.3), reducing screening workload by 48.3%. Model-based full-text screening achieved sensitivity of 82.6% (95% CI: 66.7-95.8) and AUC of 83.2% (95% CI: 74.5-91.9), reducing workload by 81.3%.
    CONCLUSION: A machine-assisted screening workflow reduced workload. The workflow extended workload reduction to full-text screening and embedded community researchers throughout, enabling the completion of a scoping review with nuanced eligibility criteria.
    PLAIN LANGUAGE SUMMARY: Every year, more scientific papers are published. This makes it hard for researchers to keep up with all the information when they are doing studies that summarize what we know; these studies are called reviews. One of the most time-consuming parts of a review is screening thousands of papers for fit to the review topic. Researchers usually begin with title and abstract screening, which means reading the title and a short summary of each paper (called an abstract) to narrow down, before full-text screening, where they read entire papers to make final inclusion decisions. Artificial intelligence (AI) tools such as machine learning models can help speed up the review process by screening papers. However, many AI tools are developed without knowledge user input, which can improve the relevance of review outcomes and uncover nuances in the review topic, such as assessing the engagement of lay communities in research, which is often inconsistently defined and applied. In this study, we implemented an AI-assisted process to support title and abstract, and full-text screening. We applied the process to a review of approaches used to engage lay communities in modeling studies of infectious diseases. The study was led by knowledge users from community-based organizations in collaboration with academic partners, who together defined the review topic and screened titles and abstracts and full texts. First, we trained a machine learning model on a subset of titles and abstracts that were screened by human reviewers. Then, for full-text screening, we used a two-step approach: we first narrowed down full texts by identifying those that were more likely to fit the review topic based on the words and phrases they used, and then applied a second machine learning model trained on full texts screened by human reviewers. The AI-assisted process helped us screen more than 9,000 papers, reducing screening workload by 48% during title and abstract screening and by 81% during full-text screening. These findings show that combining AI tools with knowledge user input can support reductions in screening workload while accounting for nuanced and inconsistently reported information.
    Keywords:  full-text; human-in-the-loop; machine learning; scoping review; screening; title and abstract
    DOI:  https://doi.org/10.1016/j.jclinepi.2026.112541
  2. Cancer Innov. 2026 Oct;5(5): e70077
       Background: Systematic reviews (SRs), often conducted for evidence synthesis, require extensive multistage screening, with Stage-I title/abstract screening being especially time consuming. Covidence systematic review software, a widely used SR platform, has offered artificial intelligence (AI)-assisted Stage-I screening function since 2022. However, this function's performance has not yet been formally compared across heterogeneous, large-scale oncology-focused SRs.
    Methods: We selected two completed SRs conducted by the Program in Evidence-based Care (PEBC), Ontario, Canada. The first SR (Axilla-BC, N = 8774) supported a clinical practice guideline in breast cancer and the second SR (PET-utility, N = 7267) informed the provincial Positron Emission Tomography (PET) steering committee in Ontario. Using each SR, we conducted a simulation-based study by drawing subsets of 500, 1000, and 2000 articles and running 30, 30, and 10 simulation trials, respectively, to emulate Covidence's AI-assisted Stage-I screening. For each trial, we calculated the workload and time savings at 95% and 100% sensitivity for relevant articles and at 100% sensitivity for finally included articles as primary outcomes. Secondary outcomes comprised missed finally included articles at 95% sensitivity for relevant articles. Wilcoxon rank-sum tests were used to compare outcomes between these two SRs.
    Results: When pooling all subset sizes and the N = 1000 subsets, workload savings at 95% sensitivity were significantly higher for Axilla-BC than for PET-utility (median = 37.4% vs. 26.2%, W = 3156.5, p = 0.003 and median = 38.3% vs. 20.5%, W = 641.0, p = 0.005, respectively). Corresponding time savings were also significantly higher for Axilla-BC (median = 6.0 vs. 4.2 h, W = 3089.5, p = 0.008 and median = 9.6 vs. 5.1 h, W = 639.0, p = 0.005). In Axilla-BC, 20% of the 70 trials missed a total of 18 finally included articles, but only one PET-utility trial missed one article.
    Conclusions: In two large oncology-SRs, Covidence's AI-assisted Stage-I screening function demonstrated meaningful potential to reduce reviewer workload and screening time. Nonetheless, Covidence's performance varied notably between reviews, suggesting a possible association between the results and the characteristics of the SR, including topic, methodological complexity, relevance rate, and inclusion rate. Careful human oversight remains essential to ensure that no potentially finally included studies are missed in evidence synthesis, particularly in nuanced or methodologically complex reviews.
    DOI:  https://doi.org/10.1002/cai2.70077
  3. J Clin Epidemiol. 2026 Sep 28. pii: S0895-4356(26)00418-X. [Epub ahead of print] 112542
       OBJECTIVE: We aimed to determine the accuracy of one versus two reviewers for full-text screening, and serial screening (title/abstract and full text). We also determined the inter-rater reliability, post-test probability (PTP), false positive and negative rates, and sources of misclassification.
    DESIGN AND SETTING: One reviewer screened fulltext articles from three systematic reviews evaluating interventions (acupuncture, education and TENS) for chronic low back pain. We combined these results with titles and abstracts screened by the same reviewer to assess serial screening accuracy (screening title and abstracts then full text studies). Outcomes were compared to consensus decisions from two independent reviewers. We computed sensitivity, specificity, positive (PPV) and negative predictive values (NPV) with 95% confidence intervals (CIs). Inter-rater reliability (Cohen's kappa) and positive PTP for serial screening were calculated. Rates and reasons for misclassification were examined. A sensitivity analysis of English studies explored the impact of AI-assisted translation of non-English studies.
    RESULTS: A total of 120, 92 and 91 studies were screened for the acupuncture, education, and TENS reviews, respectively. Compared with two reviewers, the sensitivity of single reviewer full-text screening ranged from 62% to 88%, specificity from 84% to 96%, PPV from 70% to 77% and NPV from 89% to 97%. Kappa ranged from 0.61 to 0.73. Single reviewer sensitivity for serial screening ranged from 62% to 81%, specificity from 98% to 100%, PPV from 68% to 71% and NPV was 99%. Kappa ranged from 0.67 to 0.73. Positive PTP ranged from 42.7% to 63.5%. Misclassifications were mostly due to interpretation of the population, interventions or comparators. There were no or small differences in the estimates with overlap of CIs between English and all language analyses.
    CONCLUSIONS: The accuracy of one versus two reviewers may be reduced in serial screening, resulting in missed and falsely included studies, ranging across reviews. Therefore, there is uncertainty around whether a single reviewer can correctly select relevant studies. However, the impact on rapid review results is unknown. Although overall agreement is substantial, reporting sensitivity and specificity can provide insight into classifications between reviewers. Carefully defining the population, intervention and comparators in screening is required and highlights the need for improved reporting in trials.
    Keywords:  accuracy; meta-research; rapid review; reliability; screening; systematic review
    DOI:  https://doi.org/10.1016/j.jclinepi.2026.112542
  4. Front Res Metr Anal. 2026 ;11 1882130
       Introduction: Large language models (LLMs) have been widely applied to text classification; however, their effectiveness in domain-specific scholarly metadata classification remains insufficiently explored. This study investigates the use of open-source LLMs to classify modeling and simulation research articles across three predefined metadata dimensions: Industry, Application Area, and Modeling Method.
    Methods: Using a curated dataset of research articles from the AnyLogic research repository, we evaluated models from the Qwen2.5, Llama3.1, Mistral, Ministral, and Gemma3 families. We compared Direct, Few-shot, and Chain-of-Thought prompting using either the article title alone or the title and abstract together. Performance was evaluated using Macro-F1 across the three metadata dimensions.
    Results: The best overall configuration, Ministral-3-14B with Direct prompting and the title and abstract as input, achieved a Macro-F1 of 53.8%. Performance varied substantially across the metadata dimensions. Modeling Method was the easiest to classify, reaching a maximum Macro-F1 of 69.1% with Ministral-3-14B. The best Macro-F1 scores for Industry and Application Area were 57.6% and 39.4%, respectively. Including abstracts consistently improved classification performance across the evaluated models, whereas differences among the three prompting strategies were relatively small. Categories that were semantically overlapping, broadly defined, highly imbalanced, or not explicitly mentioned in the article text were particularly difficult to classify.
    Discussion: The findings indicate that open-source LLMs can support domain-specific scholarly metadata classification without task-specific fine-tuning. However, their moderate and dimension-dependent performance limits their suitability for fully automated fine-grained metadata enrichment. Richer contextual information, more clearly defined taxonomies, and human validation are therefore needed to improve the reliability of LLM-assisted scholarly metadata classification.
    Keywords:  large language models; prompt engineering; scholarly metadata classification; simulation and modeling research; text classification
    DOI:  https://doi.org/10.3389/frma.2026.1882130
  5. BMJ Open. 2026 Sep 30. 16(9): e123460
      Systematic reviews are foundational to evidence-based practice, but the reproducibility and auditability of study identification and selection remain persistent challenges. Existing reporting standards, such as Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 and PRISMA-Search extension, have improved transparency, yet important gaps remain in documenting search strategies, retrieved records, screening decisions and supplementary identification methods. In this article, we discuss key barriers to reproducibility, including incomplete or inconsistent search documentation, limitations in date filter precision, restricted access to databases and platforms, non-repeatable supplementary search methods and variability in applying eligibility criteria. We suggest a minimum documentation set specifying the search histories, source-specific exports, record-management information, screening decisions and artificial intelligence (AI)-related records that review teams should preserve and share. We also recommend information-specialist involvement and carefully documented human-in-the-loop AI use. These strategies may not guarantee perfect reproducibility, but they can improve auditability, reduce research waste and strengthen trust in systematic reviews.
    Keywords:  Artificial Intelligence; Literature; Meta-Analysis; STATISTICS & RESEARCH METHODS; Systematic Review
    DOI:  https://doi.org/10.1136/bmjopen-2026-123460
  6. medRxiv. 2026 Sep 23. pii: 2026.09.17.26363288. [Epub ahead of print]
       Background: Clinicians increasingly use AI systems to search the medical literature, but current benchmarks do not jointly test whether a response identifies the originating paper and its supporting passage. LitBench evaluates four dimensions: stated versus interpretation-requiring facts, the number and composition of competing papers, one-versus two-paper evidence, and refusal when evidence is absent.
    Methods: LitBench comprises 1,980 open-access papers (1,000 epilepsy papers holding the answers, 980 decoys from unrelated fields) and 9,472 human-reviewed facts, giving 2,188 single-paper questions and 170 requiring a fact from each of two papers. Difficulty rose by burying the answer among up to 2,000 distractors, then again over live PubMed Central. The same questions were then asked with the answering paper removed, so that refusal was the correct response. Four systems were tested: Gemma-4B, Gemma-12B, Sonnet-5, and DeepSeek-V4-Flash (refusal conditions only). Three model judges scored each answer by majority; two-paper questions counted only when both facts were found.
    Results: Across fixed-corpus single-paper conditions, accuracy ranged from 55.9% to 71.9% for Gemma-4B, 61.2% to 75.5% for Gemma-12B, and 90.0% to 91.3% for Sonnet-5. Similar epilepsy papers were harder than mixed candidate sets for both Gemma configurations. Two-paper accuracy was 5.7%, 17.1%, and 37.8%, respectively. With evidence absent, Gemma-12B refusal fell from 88.3% to 39.9% as candidate sets grew; Gemma-4B almost never refused. Sonnet-5 refused on 94.9% of sampled single-paper and all sampled two-paper cells. DeepSeek-V4-Flash also refused frequently, but on 80.3% of matched answer-present controls.
    Conclusions: No system was reliable across every axis tested: accuracy fell as competing papers were added, interpretation-requiring facts were harder than stated ones, two-paper synthesis was rarely achieved, and only some systems recognized absent evidence. What counts as "successful" literature search carries considerable nuance; LitBench can evaluate proposed tools at a fine-grained level.
    DOI:  https://doi.org/10.64898/2026.09.17.26363288
  7. medRxiv. 2026 Sep 21. pii: 2026.09.15.26362921. [Epub ahead of print]
      Medical evidence synthesis increasingly requires published studies, real-world clinical data and structured biomedical knowledge, yet most automated systems remain centered on literature retrieval and summarization. Here we present MES, a multi-agent framework for source-grounded medical evidence synthesis. MES coordinates six specialized agents to decompose clinical questions, retrieve literature and trial evidence, generate question-specific RWE by constructing and analyzing real-world cohorts, query biomedical knowledge graphs, integrate quantitative findings and screen generated claims against their cited evidence. The framework uses evidence-based medicine taxonomies to clarify underspecified questions, adapts evidence use when sources are absent or discordant and reports unresolved gaps rather than forcing consensus. MES also provides claim-level provenance tracking and cross-agent consistency checking, allowing final reports to distinguish trial evidence, real-world associations and knowledge-graph support. We evaluated MES in two clinical use cases and across 144 clinical queries spanning six evidence-based medicine categories. MES generated structured reports that preserved source traceability, identified cross-source disagreement and communicated uncertainty across diverse clinical question types. MES provides an auditable framework for organizing, testing and contextualizing heterogeneous clinical evidence while keeping the evidentiary basis of each conclusion explicit.
    DOI:  https://doi.org/10.64898/2026.09.15.26362921
  8. JAMIA Open. 2026 Oct;9(5): ooag194
       Objectives: Exploratory data analysis (EDA) is foundational to clinical research yet inaccessible to clinician-scientists without programming training. Existing large language model (LLM) tools improve accessibility but struggle with hallucinations. We developed and evaluated data reflector (DR), a hybrid agentic platform pairing natural language interaction with deterministic statistical computation.
    Materials and Methods: Data reflector restricts the LLM to intent interpretation, tool selection, and filter parsing while routing all numerical computation through predefined deterministic functions executed locally. We piloted DR with 6 participants of diverse technical backgrounds across 6 clinical (MIMIC-IV, eICU Collaborative Research Database, National Health and Nutrition Examination Survey) and nonclinical (energy, COVID-19, urban mobility) datasets, assessing usability (system usability scale, SUS), task completion time, analytical output correctness against participant-generated manual reference results, hallucination prevention, and independent reproducibility by a second operator. A head-to-head comparison with ChatGPT 5.3 chatbot was also conducted.
    Results: Data reflector achieved a mean SUS of 94.58, with reduced task completion time across all evaluable participants. Statistical outputs and generated distribution curves demonstrated complete agreement with participant-generated manual reference results across all evaluable outputs. Independent reexecution by a different operator using independently constructed prompts also demonstrated complete agreement with the original DR outputs across all evaluated tasks. Hallucination-prevention safeguards correctly declined out-of-scope and impossible queries. ChatGPT produced incomplete, internally inconsistent outputs on identical tasks where DR returned complete, accurate results.
    Discussion: Data reflector's architectural separation of LLM reasoning from numerical computation removes dependence on prompt phrasing, session history, and stochastic LLM sampling while preserving conversational accessibility, addressing reproducibility limitations of general-purpose LLM tools.
    Conclusion: Data reflector offers a practical pattern for reproducible, accessible artificial intelligence-assisted EDA across heterogeneous data sources in modern clinical and translational research.
    Keywords:  agentic systems; exploratory data analysis; large language models; statistics
    DOI:  https://doi.org/10.1093/jamiaopen/ooag194
  9. Transl Behav Med. 2026 Jan 07. pii: ibag070. [Epub ahead of print]16(1):
       BACKGROUND: Artificial intelligence (AI) holds considerable potential for addressing challenges in knowledge translation and implementation research, including increasing the speed of knowledge synthesis, tailoring of implementation strategies, and collection of implementation data. To date, no prior review has systematically examined how AI has been applied within implementation research.
    AIMS: This scoping review aimed to synthesize the current evidence on applying AI in implementation research, identify reported outcomes and challenges, including ethical, regulatory, and practical considerations.
    METHODS: This review is reported following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for scoping review checklist. A systematic search was conducted across MEDLINE Complete, Scopus, IEEE Xplore, and ACM Digital Library on 31 July 2025. Seven peer-reviewed empirical studies and methodology papers were included and narratively synthesized.
    RESULTS: Included studies applied machine learning, natural language processing, and generative AI to support implementation monitoring, strategy selection, literature synthesis, and knowledge dissemination. These applications may help challenges related to speed, analytic burden, sustainability, and contextual tailoring in implementation research. Ethical considerations included privacy, transparency, and equity.
    CONCLUSIONS: The current evidence base reflects an early stage of AI application in implementation research. While offering important insights, it highlights opportunities for deeper theoretical integration and more robust empirical evaluation to advance translation of research into practice.
    Keywords:  implementation science; implementation strategy; knowledge translation
    DOI:  https://doi.org/10.1093/tbm/ibag070
  10. J Am Coll Clin Pharm. 2026 Oct;9(10): e70300
      Artificial intelligence (AI) is increasingly shaping the pharmaceutical industry. This ACCP commentary examines the implications of AI for industry-based clinical pharmacists across the pharmaceutical life cycle, including drug development, regulatory affairs, medical affairs, health economics and outcomes research, and pharmacovigilance. Artificial intelligence-enabled tools may support target identification, clinical trial design, regulatory intelligence, evidence synthesis, medical content generation, real-world evidence analysis, economic modeling, adverse event processing, and safety signal detection. As these tools mature, the role of the clinical pharmacist is likely to shift from primarily task execution toward clinical interpretation, quality assurance, strategic decision-making, and governance of AI-supported outputs. However, AI implementation also introduces important risks and practical implementation challenges. These limitations reinforce the need for clinical pharmacists to remain actively engaged as human-in-the-loop experts who can assess whether AI-generated insights are scientifically valid, clinically relevant, ethically sound, and appropriate for decision-making. Ultimately, AI may expand the reach and efficiency of pharmaceutical industry functions, but successful integration will depend on pharmacist leadership in evaluation, oversight, and responsible implementation.
    Keywords:  artificial intelligence; clinical pharmacist; medical affairs; pharmaceutical industry; pharmacovigilance; regulatory affairs
    DOI:  https://doi.org/10.1002/jac5.70300
  11. J Med Internet Res. 2026 Sep 30. 28 e90872
       BACKGROUND: Understanding patients' experiences is essential for advancing patient-centered care, especially in chronic diseases that require ongoing communication. Qualitative thematic analysis is widely used to explore these experiences; however, the process remains labor-intensive, subjective, and difficult to scale.
    OBJECTIVE: This study aimed to develop and evaluate Collaborative Theme Identification Agents (CoTI), a multiagent large language model framework designed to support manual thematic analysis by rapidly generating supporting excerpts, initial codes, and themes.
    METHODS: CoTI consists of 3 agents: Instructor, Thematizer, and CodebookGenerator. The Instructor refines instruction prompts, the Thematizer extracts supporting excerpts and generates initial codes for each transcript, and the CodebookGenerator groups similar codes across all transcripts into a codebook with themes. We evaluated CoTI primarily using 12 transcripts of patient with heart failure, with a focus on perceptions of medication intensity. CoTI-generated outputs were compared against the reference standard developed by senior investigators. To explore human-AI interaction in thematic analyses, we further implemented CoTI in a user-facing application.
    RESULTS: CoTI generated supporting excerpts, initial codes, and themes that were more similar to those of senior investigators than were the outputs of junior investigators, baseline natural language processing models, and other basic large language models. In an exploratory human-AI collaboration experiment, we found that the collaboration between CoTI and junior investigators provided only marginal gains compared to CoTI alone. A possible hypothesis was that junior investigators may overrely on CoTI and limit their independent critical thinking.
    CONCLUSIONS: CoTI can improve the efficiency of thematic analysis by rapidly generating supporting excerpts, initial codes, and themes for human researchers' review. These findings highlight CoTI's potential as a useful tool for scalable qualitative research.
    Keywords:  human-AI collaboration; large language model; multiagent framework; qualitative research; thematic analysis
    DOI:  https://doi.org/10.2196/90872