bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–08–09
eleven papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. Smart Med. 2026 Aug;5(4): e70045
      Meta-analysis is fundamental to evidence-based medicine, yet traditional workflows remain labor-intensive and susceptible to bias. Although LLM-based research agents offer opportunities for workflow automation, they often lack the data fidelity and methodological traceability required for rigorous quantitative evidence synthesis, particularly when parsing multimodal scientific charts. To address this challenge, we introduce MacAma, a semi-automated multi-agent framework for protocol-constrained and human-verifiable meta-analysis. MacAma operationalizes selected PRISMA 2020 reporting items, PICOS-based eligibility logic, and SYRCLE risk-of-bias domains as structured prompts, decision rules, output fields, and audit records. Critically, MacAma adopts a risk-aware automation strategy: Lower risk, repetitive, and protocol-driven tasks, such as literature screening and drafting, are delegated to AI agents, whereas high-impact steps that directly affect effect-size estimation and statistical conclusions, such as quantitative chart-data extraction, remain subject to expert verification. In a preclinical radiotherapy case study evaluating tumor-related immune outcomes and metastatic potential mediated by circulating tumor cells, MacAma achieved competitive screening performance in the evaluated benchmark and reduced the manual screening burden by over 80% within the current workflow. The case study further demonstrates how structured agent outputs, predefined criteria, and audit records can support transparent screening, data extraction, statistical synthesis, and manuscript drafting. These results suggest that MacAma may provide a scalable and auditable framework for AI-assisted meta-analysis, although important limitations remain in full-text access, quantitative chart data extraction, and expert interpretation of heterogeneity. MacAma is open-source and available at https://github.com/YilinYuan/MacAma.
    Keywords:  large language models; literature screening; meta‐analysis; prompt engineering; research automation
    DOI:  https://doi.org/10.1002/smmd.70045
  2. Discov Public Health. 2026 ;23(1): 1260
      Habitual physical activity (PA) is a fundamental determinant of health, yet its assessment and reporting vary substantially across research studies, creating persistent challenges for evidence synthesis. Systematic reviews that aim to synthesise habitual PA outcomes are frequently hindered by heterogeneity in measurement tools, implementation protocols, outcome metrics, and reporting practices. These heterogeneities complicate decision-making regarding study inclusion, data extraction, data analysis, reduce the reliability of meta-analysis, and limit transparency, comparability, and interpretability of systematic review findings. Drawing on insights from a recent systematic review, this commentary highlights key methodological challenges specific to assessment and reporting of habitual PA within systematic reviews. We argue that the absence of standardised criteria for identifying, screening, and appraising habitual PA measures is a major contributor to these challenges, and undermines transparency, comparability, and interpretability of systematic review findings. To address this gap, we present a protocol for a structured screening guide, describe its conceptual domains, and discuss the need for empirical validation. We call for a collaborative, consensus-driven process to test its reliability, feasibility, and validity before wider adoption. Establishing agreed methodological guidance is essential to strengthen the quality, credibility, and policy relevance of systematic reviews of habitual PA.
    Keywords:  Habitual physical activity; Measurement heterogeneity; Methodological guidance; Physical activity assessment; Systematic reviews
    DOI:  https://doi.org/10.1186/s12982-026-02612-8
  3. PLoS One. 2026 ;21(8): e0351135
       BACKGROUND: Each day, over 100 randomized controlled trials (RCTs) are published, making it impossible for clinicians to stay up-to-date with medical literature. Large language models (LLMs) can identify and summarize emerging clinical evidence and support medical education.
    METHODS: We created and prospectively evaluated a newsletter, Trial Files, which leverages an LLM to summarize RCT abstracts relevant to general internal medicine. We created a software tool, called PaperScrape, which leverages the Medline application programming interface (API) to identify trials published in five high-impact journals. Information from each RCT's abstract was extracted, and plain-language summaries were generated using OpenAI's LLM API. We analyzed the accuracy of summaries generated by an LLM (compared to manual review), results of a subscriber survey, and effectiveness of marketing strategies on user growth.
    RESULTS: From June 2023 to March 2025, 50 newsletters with 3 RCTs each were distributed to 648 subscribers. A subset of 96 RCTs was randomly selected to evaluate reporting accuracy with prompt engineering. The accuracy for reporting study information with prompt engineering, compared to manual review, was 97.1% for study phase, 92.2% for blinding, 85.4% for sample size, 97.9% for patient population, 94.7% for comparison groups, and 92.7% for primary outcome. Forty-three subscribers completed a survey about Trial Files. The mean overall rating was 4.7 out of 5 (5 representing "very good"), and all respondents agreed the newsletter made it easier to keep up-to-date with emerging clinical trials in internal medicine. The most effective strategy for user growth was promotion at a meeting, conference, or education session (6.8 subscribers per day, compared to 0.7 subscribers gained per day on days without promotion, p < 0.0001).
    CONCLUSION: LLMs can provide concise, accurate summaries of RCTs, which can help general internists stay up-to-date on recently published trials.
    DOI:  https://doi.org/10.1371/journal.pone.0351135
  4. Nat Mach Intell. 2026 Jul;8(7): 1142-1156
      Compared with generic artificial intelligence agents, deep research agents perform longer-horizon reasoning and deeper literature exploration to investigate complex questions. Here we present DeepEvidence, a deep research agent for evidence exploration and synthesis across heterogeneous biomedical knowledge sources. DeepEvidence advances deep research through coordinated multi-agent collaboration combining breadth-first and depth-first research strategies to search, explore and aggregate evidence from multiple biomedical knowledge bases and literature. It also incrementally constructs an evidence graph of key entities and observations to support transparent tracking, attribution and validation of the research process. DeepEvidence substantially outperforms generic artificial intelligence agents across four open benchmarks. We further establish seven benchmark tasks spanning major stages of biomedical discovery, including drug discovery, preclinical experimentation, clinical trial development and evidence-based medicine. DeepEvidence demonstrates substantial improvements in systematic evidence exploration and synthesis. These results highlight the potential of deep research agents to accelerate biomedical discovery and translational research.
    DOI:  https://doi.org/10.1038/s42256-026-01266-0
  5. Neurosci Inform. 2026 Mar;pii: 100262. [Epub ahead of print]6(1):
      Careful evaluation of research methodology is fundamental to scientific progress but represents a significant burden on human experts. The complexity of functional MRI (fMRI) methods makes transparent reporting, as suggested by OHBM COBIDAS guidelines, particularly critical. Large Language Models (LLMs) present a potential solution for rapid, scalable methodological assessment. We evaluated three state-of-the-art LLMs (Gemini 2.5 Pro, Claude 4 Sonnet, ChatGPT-o3-pro) against human expert ratings. Fifty fMRI articles (taken from 2016 to 2025) were independently evaluated by ten human experts and three LLMs using an 82-item COBIDAS based rubric. Human raters demonstrated excellent inter-rater reliability (ICC = 0.801), while LLMs showed poor internal agreement (ICC = 0.254). When comparing total scores across papers, Gemini showed strong positive correlation with human consensus (r = 0.693, p < 0.0001), Claude showed moderate positive correlation (r = 0.394, p = 0.004), while ChatGPT showed negative correlation (r = -0.172, p = 0.233). Gemini maintained high reliability when added to human raters (combined ICC = 0.811), achieving 85.3 % exact agreement and 98.8 % within-1-point agreement. Domain-specific analysis revealed Gemini's consistently high agreement across all six COBIDAS sections (experimental design: 0.915, statistical modeling: 0.880), while ChatGPT and Claude showed weaker, more variable performance. Obvious differences emerged in determining non-applicable items: humans marked 40.5 % as not applicable versus 32.3 % for Gemini, 9.2 % for ChatGPT and 21.1 % for Claude. ChatGPT exhibited extreme score volatility, with papers ranging from 0 to 121 points compared to humans' 44.2-77.7 range. LLM scoring required 1-7 min versus 30-35 min for humans. This proof-of-concept study demonstrates that LLM-assisted methodological evaluation is feasible for complex neuroimaging research and could likely be applied to other research fields.
    DOI:  https://doi.org/10.1016/j.neuri.2026.100262
  6. J Am Med Inform Assoc. 2026 Aug 06. pii: ocag123. [Epub ahead of print]
       OBJECTIVE: To develop and systematically compare a human-led and LLM-assisted hybrid deductive-inductive workflow for qualitative analyses.
    MATERIALS AND METHODS: We analyzed 122 transcripts (n = 61 research clinical consultations; n = 61 reflexive interviews) from a video ethnography study of patients with heart failure. Human-led thematic analysis used Dedoose software, and LLM-based analysis was conducted using ChatGPT Edu (GPT-5.2; OpenAI) with an eleven-prompt protocol. Both applied a hybrid deductive-inductive approach. The research team compared outputs across 63 human-LLM theme pairs using human consensus and LLM-based evaluation, integrated themes into a final framework, and manually verified quotation fidelity against the original transcripts.
    RESULTS: Human-led and LLM-generated analyses produced complementary cross-cutting themes (7 human; 9 LLM), all judged valid and integrated into 14 final themes across three domains. Thematic overlap was moderate to substantial (Hit Rate 1.00; Jaccard 0.44-0.51). Robustness testing across three runs revealed recurrence of five core concepts alongside variability in theme labels and counts. Quotation fidelity showed 68% verbatim, 20% paraphrased, 6% partial and 3% full hallucinations, and 3% truncated excerpts; verbatim quotations did not always clearly support their assigned themes.
    DISCUSSION: LLM-assisted analysis is feasible for large-scale qualitative health research within a HIPAA-compliant environment using an adaptable eleven-prompt protocol. Human oversight remained essential for contextual interpretation, quotation verification, and assessment of theme-quotation support.
    CONCLUSIONS: LLMs are best positioned as analytic partners rather than autonomous coders. Transparent workflows with human-in-the-loop validation are essential for responsible AI integration in health and biomedical informatics.
    Keywords:  Large Language Models (LLMs); clinical informatics; patient-reported outcomes; qualitative research; thematic analysis
    DOI:  https://doi.org/10.1093/jamia/ocag123
  7. Evid Based Toxicol. 2025 ;3(1): 2485111
      The environmental health vocabulary (EHV) represents manually curated terminologies developed by the US Environmental Protection Agency's (EPA) Chemical Pollutant Assessment Division (CPAD) for standardizing reporting of health effect information. Recognizing that manual data curation is a resource bottleneck, a semi-automated curation workflow was realized. The objectives of this work are to describe the manual creation of the EHV and improve the efficiency of manual data curation by implementing a new semi-automated curation workflow that minimizes manual review using a sequence of computational text analysis and quality assurance/quality control (QA/QC) steps with a high level of accuracy. To facilitate semi-automated curation a sequence of computational text analysis and manual steps were developed. Described are (1) a series of computational text processing steps to normalize and match extracted terms to the EHV, (2) a QA step of the computationally identified matches; (3) a manual review of unmatched terms; and (4) curation of the EHV that includes completion of missing hierarchical data and related metadata. The EHV was manually created to promote data aggregation, integration, accessibility and transparent data exchange across EPA partners by normalizing the data extracted into the EPA Health Assessment Workplace Collaborative (HAWC). The workflow described here removes the manual curation bottleneck by transforming data curation into a streamlined semi-automated process powered by computational text processing steps. This semi-automated curation method offers several advantages to the environmental health community including (but not limited to) efficiency by automating repeating data management tasks, scalability to a large volume of terms and terminology resources, and better integration with other data sets and artificial intelligence (AI) and machine learning (ML) models.
    Keywords:  Ontology; artificial intelligence; controlled vocabulary; data curation; health information; interoperability/standards; risk assessment; systematic review
    DOI:  https://doi.org/10.1080/2833373X.2025.2485111
  8. Eur Urol Focus. 2026 Aug 06. pii: S2405-4569(26)00121-5. [Epub ahead of print]
    Next-Gen Research Group
      This study aimed to externally validate the performance of the European Association of Urology (EAU) Guidelines Bot in neuro-urology by assessing the accuracy, completeness, and clarity of chatbot-generated answers to guideline-based questions and to compare its performance with that of a general-purpose large language model (ChatGPT 5.5). A cross-sectional validation study was conducted using 47 questions derived from the EAU Neuro-Urology Guidelines. Each question was linked to a specific recommendation and classified by recommendation strength (strong vs weak). Questions were independently submitted to both the EAU Guidelines Bot and ChatGPT 5.5 without additional prompting. Two expert urologists independently evaluated each response for accuracy, completeness, and clarity using a five-point Likert scale; discrepancies were resolved by a third reviewer. Overall, 45 questions (95.7%) were linked to strong recommendations and two (4.3%) to weak recommendations. The EAU Guidelines Bot and ChatGPT 5.5 achieved identical mean accuracy scores (4.96 ± 0.20), with all responses rated as highly accurate (Likert 4-5). ChatGPT 5.5 indicated significantly higher completeness scores than did the EAU Guidelines Bot (4.74 ± 0.44 vs 4.57 ± 0.54; p = 0.011), whereas clarity scores were not significantly different (4.83 ± 0.38 vs 4.77 ± 0.43; p = 0.083). High-quality completeness was observed in 46/47 EAU Guidelines Bot responses (97.9%) and 47/47 ChatGPT responses (100%). Score discrepancies between systems were identified in ten of 47 questions (21.3%) and were limited to completeness and clarity domains. Performance remained uniformly high across recommendation grades, with no meaningful differences observed. The EAU Guidelines Bot showed excellent accuracy, completeness, and clarity when applied to neuro-urology guideline-based questions. Its performance was comparable to that of ChatGPT 5.5, with both systems providing highly accurate guideline-concordant responses. Although ChatGPT 5.5 generated more comprehensive answers, the EAU Guidelines Bot maintained closer adherence to the original guideline recommendations. Although not a substitute for clinical judgment, the tool appears to be a reliable adjunct for rapid access to evidence-based neuro-urological guidance.
    Keywords:  Artificial intelligence; Clinical guidelines; Decision support systems; Neurogenic lower urinary tract dysfunction
    DOI:  https://doi.org/10.1016/j.euf.2026.06.025
  9. Int J Impot Res. 2026 Aug 07.
    EAU Young Academic Urologists (YAU) Sexual and Reproductive Health Working Group
      Artificial intelligence (AI)-based language models are increasingly explored as tools for interpreting and applying clinical guideline recommendations. In urology, the European Association of Urology (EAU) recently introduced a guideline-specific chatbot; however, its comparative performance relative to contemporary general-purpose large language models (LLMs) remains unclear. In this structured comparative study, five AI systems-the EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot - Smart GPT-5, and Perplexity Pro-were evaluated using 13 clinical questions derived directly from strongly recommended statements in the EAU erectile dysfunction (ED) guidelines. Responses were independently assessed by three senior reviewers across five predefined domains: relevance, clarity, structure, clinical utility, and factual accuracy, using a 5-point Likert scale. The primary outcome of the study was the composite performance score, which was calculated as the mean of the five domain scores. Inter-rater reliability was calculated using ICC(2,k), and differences among models were analyzed with the Friedman test followed by Holm-adjusted Wilcoxon post-hoc comparisons. Significant performance differences were observed across all domains (all p < 0.001). The highest composite scores were observed for Gemini 2.5 Pro [4.60 (4.40-4.73)] and the EAU Guidelines Bot [4.53 (4.47-4.80)], followed by ChatGPT-5 [4.27 (4.07-4.47)]. Lower composite scores were observed for Copilot - Smart GPT-5 [3.73 (3.40-3.87)] and Perplexity Pro [3.60 (3.47-3.80)]. Domain-level analysis showed consistently high median scores (≥ 4) for factual accuracy among top-performing models, whereas variability was more pronounced in clarity, structure, and clinical utility. These findings suggest that both guideline-specific systems and advanced general-purpose LLMs may generate responses broadly consistent with guideline-based recommendations in structured ED scenarios. However, variability across domains-particularly in structure and clinical utility-and modest differences in composite performance suggest that these models should be interpreted as supportive tools rather than definitive clinical decision-making systems, requiring further validation in real-world settings.
    DOI:  https://doi.org/10.1038/s41443-026-01339-z
  10. Front Artif Intell. 2026 ;9 1908389
      
    Keywords:  agentic AI; artificial intelligence; interdisciplinary research; laboratory automation; large language models; machine learning; responsible AI; scientific discovery
    DOI:  https://doi.org/10.3389/frai.2026.1908389