bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–08–30
twenty-two papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. J Healthc Inform Res. 2026 Sep;10(3): 654-665
      Systematic reviews are a key component of evidence-based medicine, playing a critical role in synthesizing existing research evidence and guiding clinical decisions. However, with the rapid growth of research publications, conducting systematic reviews has become increasingly burdensome, with title and abstract screening being one of the most time-consuming and resource-intensive steps. To mitigate this issue, we designed a two-stage dynamic few-shot learning (DFSL) approach aimed at improving the efficiency and performance of large language models (LLMs) in the title and abstract screening task. Specifically, this approach first uses a low-cost LLM for initial screening, then re-evaluates low-confidence instances using a high-performance LLM, thereby enhancing screening performance while controlling computational costs. We evaluated this approach across 10 systematic reviews, and the results demonstrate its strong generalizability and cost-effectiveness, with potential to reduce manual screening burden and accelerate the systematic review process in practical applications.
    Supplementary Information: The online version contains supplementary material available at https://doi.org/10.1007/s41666-026-00246-8.
    Keywords:  Evidence Synthesis; Large Language Models; Literature Screening; Systematic Review
    DOI:  https://doi.org/10.1007/s41666-026-00246-8
  2. Cureus. 2026 Jul;18(7): e113428
      Artificial intelligence (AI) holds great promise for improving systematic reviews, particularly as the number of reviews keeps growing. A solution to keep systematic reviews up to date are the so-called living systematic reviews. Although they could be a solution, they currently face challenges regarding everything having to be done manually. A potential solution could be AI-driven reviews that automate each step, from literature searching and translation to data extraction and analysis. Potentially, AI could search through large amounts of studies, sort them by relevance, and extract relevant data. This significantly reduces the workload for researchers, and thereby gives researchers time to focus on oversight rather than manual tasks. However, these methods raise concerns about algorithmic bias and transparency. Proper training, clear ethical guidelines, and interdisciplinary collaboration are keys to ensuring quality and integrity. AI-driven reviews may become essential for efficiently handling the expanding scientific literature.
    Keywords:  artificial intelligence in medicine; living systematic reviews; medical education; research methodology and ethics; systematic reviews and meta-analyses
    DOI:  https://doi.org/10.7759/cureus.113428
  3. PLOS Digit Health. 2026 Aug;5(8): e0001668
      Systematic reviews (SRs) are key to evidence-based medicine but are often labor-intensive, especially in the study selection step. This study assessed the use of large language models (LLMs) to automate SR study screening and selection. Five SR projects were included: two published therapeutic SRs (SR1-2), two ongoing emulated-trial SRs (SR3-4), and one economic evaluation SR (SR5). The total number of studies screened for each SR was 3,966, 3,147, 695, 3,096 and 485, respectively, with 20, 24, 46, 32 and 70 eligible studies. Three LLMs-Gemini 2.0 Flash, Llama 3.1, and Qwen 2.5-were evaluated using training sets (five studies), title/abstract datasets, and full-text datasets predicted as relevant. Prompts based on the PICOS framework were iteratively refined using a recall-first strategy. Outputs were compared with human reviewer classifications using recall, number needed to screen (NNS), and percentage reduced workload with 95% confidence intervals. In the title/abstract screening phase, Llama 3.1 and Gemini 2.0 Flash achieved consistently high recall (90.00%-100.00% and 90.48%-100.00%), with workload reduction of 61.92%-97.10% and 65.21%-97.03%, respectively. Qwen 2.5 achieved the highest workload reduction (76.16%-99.19%) but showed the lowest recall (76.67%-88.89%). In the full-text selection phase, Llama 3.1 achieved the highest recall (93.33%-100.00%) with workload reductions of 73.15%-97.41%, but slower processing time (approximately 2.3-3.6 minutes per document). Qwen 2.5 yielded lower recall (66.67%-89.71%), despite the highest workload reduction (80.82%-99.44%) and similarly slow inference times (approximately 3.0-4.2 minutes per document). Gemini 2.0 Flash balanced high recall (83.33%-100.00%) with substantial workload reduction (76.71%-98.91%) and markedly faster inference (approximately 4-8 seconds per document). LLMs-particularly Llama 3.1 and Gemini 2.0 Flash-can substantially reduce SR screening workload while maintaining high recall when guided by a recall-first prompting framework. Remaining challenges include reproducibility in closed-source models and generalizability across diverse SR topics.
    DOI:  https://doi.org/10.1371/journal.pdig.0001668
  4. BMC Nurs. 2026 Jul 29. pii: 763. [Epub ahead of print]25(1):
       BACKGROUND: The rapid development of large language models has accelerated the use of artificial intelligence (AI) in qualitative research. AI-assisted thematic analysis may improve coding efficiency and data organization, but concerns remain about its ability to interpret emotionally nuanced, contextually embedded, and culturally situated nursing narratives. Existing studies have emphasized coding performance and efficiency, whereas ethical and interpretive concerns in nursing qualitative research remain insufficiently examined.
    OBJECTIVE: This secondary qualitative analysis compared human-led reflexive thematic analysis and ChatGPT-assisted thematic analysis of the same nursing interview dataset and explored the advantages, limitations, and ethical concerns of AI-assisted qualitative inquiry in nursing.
    METHODS: Semi-structured interview data originally collected within a broader nursing research platform project were analyzed through two separate pathways: human-led reflexive thematic analysis based on Braun and Clarke's approach and ChatGPT-assisted thematic analysis. The research team compared the outputs to examine convergence, divergence, contextual simplification, thematic representation, and ethical implications. Ethical concerns were interpreted through comparative and reflexive analysis.
    RESULTS: The research team observed that AI-assisted analysis supported rapid data organization, preliminary coding, and identification of explicit thematic patterns. However, the two analytic pathways differed in their representation of contextual meanings, tacit nursing knowledge, and experience-based nursing narratives. Through comparative and reflexive interpretation, four ethical themes were identified: privacy and data security concerns; loss of contextual and emotional understanding; algorithmic reductionism and cultural bias; and changing roles and responsibilities in nursing research. AI-assisted outputs tended to be semantically coherent but contextually simplified.
    CONCLUSIONS: In this study, AI-assisted qualitative analysis was useful as a supportive tool for preliminary thematic exploration and data organization. However, human reflexive engagement remained essential for contextual interpretation, ethical judgment, and the representation of nursing experiences. The findings suggest that ethical concerns extend beyond technical performance to contextual representation, interpretive integrity, and researcher responsibility. Future research should develop reflexive, ethically grounded human-AI collaborative approaches to qualitative nursing inquiry.
    CLINICAL TRIAL REGISTRATION: Not applicable.
    Keywords:  Artificial intelligence; ChatGPT; Ethics; Human–AI collaboration; Large language models; Nursing research; Qualitative research; Reflexive thematic analysis; Thematic analysis
    DOI:  https://doi.org/10.1186/s12912-026-05085-x
  5. Eur J Phys Rehabil Med. 2026 Aug 27.
       BACKGROUND: Rehabilitation is a core component of health systems worldwide, yet its conceptual boundaries remain heterogeneous and difficult to operationalize in research and evidence synthesis. Different classification systems have been proposed, including the pragmatic framework introduced by Levack et al. in 2019 and the more structured research-oriented definition developed by Cochrane Rehabilitation in 2022. In parallel, large language models such as ChatGPT have demonstrated increasing potential in automated literature analysis. However, their ability to apply complex domain-specific classification systems has not been systematically evaluated. The aim is to compare the performance of ChatGPT-4.5, expert reviewers, and hybrid human-AI workflows in classifying research papers as rehabilitation or non-rehabilitation using two different classification systems.
    METHODS: A retrospective inter-rater reliability and diagnostic performance study was performed using online evidence synthesis and Cochrane Library. The sample consisted of 152 Cochrane systematic reviews published between November 2024 and May 2025. Each review was classified as rehabilitation or non-rehabilitation using the Levack framework and the Cochrane Rehabilitation definition. Three evaluation approaches were compared: AI-only classification (using ChatGPT), expert reviewer consensus (gold standard), and a hybrid human-AI workflow with randomized decision order (AI-first or human-first). Agreement with the expert consensus was assessed using Cohen's kappa (k). Diagnostic performance metrics and potential cognitive biases (automation and conservatism bias) in AI-assisted decisions were also analyzed.
    RESULTS: AI-only classification showed moderate agreement with expert consensus when using the Levack framework (k=0.60, 95% CI 0.47-0.71) and substantial agreement when using the Cochrane Rehabilitation definition (k=0.82, 95% CI 0.69-0.92). Hybrid workflows consistently achieved near-perfect agreement (k-range: 0.97-1.00). AI-only classification demonstrated high negative predictive values but lower positive predictive values, indicating a tendency toward overclassification of borderline cases. No evidence of automation bias or conservatism bias was observed in human-AI collaboration.
    CONCLUSIONS: Large language models demonstrate meaningful capability in classifying rehabilitation research, although performance varies with the conceptual structure of the applied classification system. Hybrid human-AI workflows achieved the highest accuracy, suggesting that AI is most effective when used as a decision-support tool rather than as an autonomous classifier. Integrating hybrid human-AI workflows into evidence synthesis can significantly accelerate the identification and mapping of rehabilitation literature, ensuring high methodological accuracy without introducing cognitive automation biases for researchers.
    DOI:  https://doi.org/10.23736/S1973-9087.26.09664-4
  6. Cochrane Database Syst Rev. 2026 Aug 27. 8 ED000179
      
    Keywords:  Evidence synthesis
    DOI:  https://doi.org/10.1002/14651858.ED000179
  7. J Manag Care Spec Pharm. 2026 Sep;32(9): 1062-1075
       BACKGROUND: Real-world evidence (RWE) is increasingly used to inform regulatory and payer policy decisions and health technology assessment, yet appraising the methodological credibility of RWE studies remains time-intensive and requires specialized expertise. The appraisal task could involve using an appraisal tool that provides a structured approach for evaluating bias in observational studies of comparative effectiveness and safety. Large language models (LLMs) may offer a scalable means to support this appraisal process, but their performance on structured bias assessment tasks has not been fully characterized.
    OBJECTIVE: To compare the performance of LLMs from 6 major artificial intelligence (AI) technology providers against human expert assessments in appraising bias in published RWE studies using the Appraisal of Potential Bias in Real-World Evidence Studies framework.
    METHODS: We conducted a comparative diagnostic accuracy study evaluating 40 LLMs from OpenAI, Anthropic, Google, xAI, Meta, and DeepSeek. Ten published RWE studies representing diverse pharmacoepidemiological designs and data sources were appraised by each LLM using a structured chain-of-thought prompt with conditional rubric injection based on the Appraisal of Potential Bias in Real-World Evidence Studies framework. Two independent human reviewers with pharmacoepidemiology training evaluated each study, with a third adjudicator resolving disagreements to establish the reference standard. LLM performance was assessed using overall accuracy and macro-averaged precision, recall, and F1 scores. Assessment time was compared between models and benchmarked against human reviewers. Bootstrap method was used to construct 95% CI for performance measures.
    RESULTS: Across 280 item-level assessments per model (10 studies × 28 items), overall accuracy ranged from 12.9% to 66.1%. The highest-performing model was Claude-Sonnet-4.6 (66.1%), followed by o3 (65.4%) and Gemini-3.1-pro-preview (65.0%). Macro-averaged F1 scores ranged from 30.9% to 66.9%; o3 achieved the highest F1 score (66.9%), followed by GROK-4 (65.5%) and Gemini-3.1-pro-preview (65.4%). Human reviewers required an average of 61.05 minutes per study; all LLMs completed assessments substantially faster, with average time per study ranging from 0.80 to 17.22 minutes relative to humans.
    CONCLUSIONS: LLMs hold considerable promise for automating methodological appraisal of RWE studies; however, their performance is variable and model dependent. Their greatest value may lie in enhancing efficiency and supporting human-led appraisal as decision-support tools rather than replacing expert review. Future research should assess performance across larger, more diverse RWE study collections and evaluate output reproducibility across repeated runs.
    DOI:  https://doi.org/10.18553/jmcp.2026.32.9.1062
  8. Surgery. 2026 Aug 01. pii: S0039-6060(26)00431-9. [Epub ahead of print]199 110506
       BACKGROUND: To examine how artificial intelligence and large language model tools may support qualitative surgical research, where they may threaten rigor and trustworthiness, and what principles should guide their responsible use.
    METHODS: This perspective reviews potential applications of artificial intelligence and large language model tools across the qualitative research process and considers their implications for analytic rigor, trustworthiness, and ethical practice.
    RESULTS: Artificial intelligence and large language model tools may be useful for bounded tasks, such as transcript summarization, data organization, preliminary code suggestion, and excerpt retrieval. However, these tools may also flatten nuance, generate hallucinated or misleading outputs, obscure interpretive processes through black-box functioning, reproduce training-data biases, and reduce researcher immersion in the data. Their responsible use depends on human oversight, verification against source data, and transparent reporting.
    CONCLUSION: Artificial intelligence and large language model tools may be most useful when applied to bounded supportive tasks rather than as substitutes for researcher interpretation, reflexivity, or analytic judgment.
    DOI:  https://doi.org/10.1016/j.surg.2026.110506
  9. Eur Heart J Digit Health. 2026 Aug;7(7): ztag125
       Aims: The rapid expansion of biomedical literature challenges clinicians' and researchers' ability to identify clinically meaningful evidence. We systematically compared five literature search tools, four artificial intelligence (AI)-assisted and one conventional, across clinically relevant cardiology research scenarios, using a blinded expert-validated gold standard to assess their ability to retrieve relevant and key references.
    Methods and results: We evaluated ChatGPT-5, Elicit, Consensus, Scite, and PubMed across four cardiology topics defined by maturity and specificity, with multiple standardized prompts. Three electrophysiology experts independently and blindly rated all retrieved references, defining two gold standards: expert-rated relevance and expert-selected key references. ChatGPT-5 achieved the highest proportion of relevant articles (90% [88-100], P < 0.001) and the highest key-reference overlap (60% [43-68], P < 0.001), whereas Scite performed lowest (20% and 10%, respectively). The tool was the primary determinant of performance (partial R 2 = 0.50), whereas prompt formulation had no significant effect. In a pre-specified subanalysis restricted to clinical studies, ChatGPT-5 and human-conducted systematic reviews overlapped by 42% (96% of shared articles highly relevant), with 58% distinct references, indicating complementary AI and human retrieval; ChatGPT-5 produced hallucinated citations when long reference lists were requested for emerging topics, underscoring the need for human verification.
    Conclusion: AI-assisted tools showed heterogeneous performance, ChatGPT-5 performing best in this cardiology setting. These preliminary, context-specific findings support hybrid human-AI strategies in which AI complements rather than replaces transparent database searches such as PubMed; larger-scale, multi-domain studies are needed to confirm and generalize them.
    Keywords:  Artificial intelligence; Cardiac resynchronization therapy; Digital health; Literature search
    DOI:  https://doi.org/10.1093/ehjdh/ztag125
  10. JMIR Form Res. 2026 Aug 25. 10 e85572
       Background: While large language model (LLM)-assisted qualitative analysis could improve the efficiency and scalability of feedback-driven curricular refinement in medical education, how best to leverage LLMs for qualitative analysis while ensuring quality outputs remains an open question. Prior work has demonstrated the feasibility of using LLMs for inductive and deductive coding tasks, but more needs to be known about how LLM-assisted thematic coding can best be deployed in a medical education context to maximize its strengths and guard against its weaknesses.
    Objective: Our study evaluated LLM performance in inductive code generation and in the deductive application of a human codebook, using a student focus-group transcript, to propose a model for AI collaboration in qualitative analysis.
    Methods: The qualitative data for this study consisted of a 1-hour focus group with 4 second-year medical students discussing a required AI-driven clinical-scenario tool (2-Sigma). Three human coders conducted an inductive thematic analysis. Using the same transcript, GPT-4o (version gpt-4o-2024-11-20; OpenAI) generated inductive codes and applied the human codebook deductively. The researchers compared the alignment between the AI inductive codes and the human consensus codebook using 3 categories: agreement, reasonable alternative, and not reasonable. Interrater reliability of AI deductive coding was evaluated using percent agreement and Cohen κ, with textual audits of discrepancies, including "misses" (failed to apply appropriate codes) and "misfires" (inappropriately applied codes). Analysis took place between February and July 2025.
    Results: In the inductive condition, GPT-4o generated 137 initial codes, of which 31.4% (n=43) demonstrated agreement with human codes, 26.3% (n=36) represented reasonable alternatives, and 42.3% (n=58) were classified as not reasonable. In the deductive condition, mean percent agreement for AI application of human codes was 96% (SD 4%, range 79%-100%) and the mean κ was 0.71 (SD 0.26, range 0-1.00). Of all 2352 coding decisions, there were 57 (2.4%) misfires and 28 (1.2%) misses; common patterns included overinterpretation of tone, failure to recognize continued ideas across excerpts, and difficulty distinguishing hypothetical vs experienced features. Based on our findings, we suggest a roadmap that retains human interpretive control while leveraging AI scalability: humans first develop a contextually grounded codebook through inductive analysis, then use AI both as a creative partner to surface alternative codes and as a tool to apply the validated codebook across the dataset.
    Conclusions: With targeted human oversight, an LLM reliably applied an existing codebook and generated additional inductive codes. These findings support a proposed workflow in which AI serves as an additional perspective within human-driven qualitative analysis, offering a scalable adjunct for qualitative analysis in medical education. Validation across larger and more diverse datasets will help confirm the generalizability of this approach.
    Keywords:  AI; generative AI; large language model; learning tool; qualitative
    DOI:  https://doi.org/10.2196/85572
  11. Bioinformatics. 2026 Aug 01. pii: btag439. [Epub ahead of print]42(Supplement_2):
    CoDiet Consortium
      Diet plays a critical role in human health, with growing evidence linking dietary habits to disease outcomes. However, extracting structured dietary knowledge from biomedical literature remains challenging due to the lack of dedicated relation extraction datasets. To address this gap, we introduce RECoDe, a novel relation extraction (RE) dataset designed specifically for diet, disease, and related biomedical entities. RECoDe captures a diverse set of relation types, including a broad spectrum of positive association patterns and explicit negative examples, with over 5000 human-annotated instances validated by up to five independent annotators. Furthermore, we benchmark various natural language processing (NLP) RE models, including BERT-based architectures and enhanced prompting techniques with locally deployed large language models (LLMs) to improve classification performance on underrepresented relation types. The best performing model was gpt-oss-20B, a locally-deployed open-weight LLM, achieving an F1-score of 61% (macro) for multi-class classification and 89% for binary classification using a hierarchical prompting strategy with a separate reflection step built in. To demonstrate the practical utility of RECoDe, we introduce the Contextual Co-occurrence Summarisation (CoCoS) framework, which aggregates sentence-level relation extractions into document-level summaries and further integrates evidence across multiple documents. CoCoS produces effect estimates consistent with established dietary knowledge, demonstrating its validity as a general framework for systematic evidence synthesis. Availability: The code, models, and dataset are publicly available at https://github.com/omicsNLP/RECoDe.
    DOI:  https://doi.org/10.1093/bioinformatics/btag439
  12. J Manag Care Spec Pharm. 2026 Sep;32(9): 1090-1100
       BACKGROUND: Health economic modeling is conceptually sophisticated but operationally repetitive and resource intensive. Recent advances in large language models suggest potential for automating components of cost-effectiveness model development.
    OBJECTIVE: To evaluate whether an agentic artificial intelligence (AI) system can reliably automate cost-effectiveness model development in the context of targeted therapies for anaplastic lymphoma kinase-positive (ALK+) non-small cell lung cancer (NSCLC).
    METHODS: We developed the Agentic Health Economic Modeling Platform (A-HEMP) to construct a cost-effectiveness model for ALK+ NSCLC therapies without access to existing models in that clinical context. Modeling decisions and extracted parameters were compared with a previously published manual cost-effectiveness analysis. A-HEMP automated PICO-based scoping, modeling approach recommendation, systematic literature review, and structured parameter extraction. Performance was evaluated across 3 domains: model structure concordance, evidence identification concordance, and parameter value alignment. Deterministic cost-effectiveness outputs were calculated externally for benchmarking.
    RESULTS: A-HEMP identified a modeling framework aligned with the published cost-effective analysis and retrieved all primary clinical trials used for clinical efficacy inputs. Concordance in model structure and assumptions was observed in 27% of modeling dimensions, with 36% partially concordant and 36% divergent. Divergences were most prominent in survival extrapolation and intracranial progression handling. Evidence identification concordance was high for primary clinical trials (100%) but moderate for cost inputs, with 25% of evidence domains fully concordant and 38% partially concordant. Parameter value alignment was high for clinical efficacy inputs and progression-free health state utilities (<2% deviation), whereas greater variability was observed for sicker health states (12% deviation) and downstream disease management costs (25%-55% deviation).
    CONCLUSIONS: Agentic AI can reliably automate upstream components of cost-effectiveness model development. However, nuanced modeling decisions with less standardized methodological guidance remain areas requiring expert oversight. These findings support a hybrid paradigm in which AI augments, but does not replace, health economists in value assessment and formulary decision support within managed care settings.
    DOI:  https://doi.org/10.18553/jmcp.2026.32.9.1090
  13. Crit Care Explor. 2026 Sep 01. 8(9): e1474
       IMPORTANCE: Large language models (LLMs) are increasingly used for scientific literature retrieval, yet their citation accuracy in specialized clinical domains remains poorly characterized. In neurocritical care (NCC), fabricated or inaccurate citations may be difficult to detect without deliberate verification.
    OBJECTIVES: To evaluate hallucination and fabrication rates of peer-reviewed citations generated by three LLMs across core NCC topics, under constrained zero-shot, memory-only conditions.
    DESIGN, SETTING, AND PARTICIPANTS: In this cross-sectional, blinded technology performance evaluation, Generative Pretrained Transformer (GPT)-5.3, DeepSeek-V3, and Grok-4 were queried on March 10, 2026, under identical zero-shot, retrieval-disabled web-interface conditions. Ten NCC topics were submitted to each model, and each model generated 10 references per topic, yielding 300 references.
    MAIN OUTCOMES AND MEASURES: Two NCC experts, blinded to model identity, independently verified each reference against PubMed, DOI, Google Scholar, and CrossRef and scored accuracy using a Hallucination Scale (0-3). The primary outcome was any hallucination, defined as any citation inaccuracy. The secondary outcome was fabrication, defined as a nonexisting complete bibliographic entity.
    RESULTS: Inter-rater agreement was excellent (κ = 0.91; 95% CI, 0.86-0.96). Overall, 165 of 300 references (55.0%) contained a citation inaccuracy, and 85 of 300 (28.3%) were completely fabricated. DeepSeek-V3 had the lowest hallucination rate (23%; fabrication 8%), followed by GPT-5.3 (69%; fabrication 27%) and Grok-4 (73%; fabrication 50%). Compared with DeepSeek-V3, Grok-4 was 3.17 times more likely to hallucinate (95% CI, 2.03-4.96; p < 0.001), and GPT-5.3 was 3.00 times more likely to hallucinate (95% CI, 1.94-4.63; p < 0.001). Topic-level findings were exploratory and should be interpreted cautiously.
    CONCLUSIONS AND RELEVANCE: Under standardized zero-shot, retrieval-disabled web-interface conditions, LLMs generated substantial numbers of inaccurate and fabricated NCC citations. Because fabricated references can appear complete and credible, artificial intelligence-generated citations should be verified across reliable databases before use in clinical, educational, or scholarly work.
    Keywords:  artificial intelligence; citation hallucination; fabricated references; large language models; neurocritical care
    DOI:  https://doi.org/10.1097/CCE.0000000000001474
  14. J Clin Epidemiol. 2026 Aug 26. pii: S0895-4356(26)00360-4. [Epub ahead of print] 112484
       BACKGROUND: Multiple stakeholders need to locate results of registered clinical trials but frequently struggle to find them. Summary results of clinical trials are often not published in trial registries, and publications containing trial results are often not explicitly linked to their respective trial registrations. Finding these results is important to researchers, systematic reviewers, research funders, regulators, clinical practitioners, and patients.
    METHODS: We developed TrialScout, a computer program that uses a large language model to match clinical trials registered on ClinicalTrials.gov with corresponding result publications indexed in PubMed. TrialScout's performance was evaluated through comparison to human-coded matches from previous studies of results reporting rates. Subsequently, TrialScout was applied in a cross-sectional analysis of a random sample of 9,600 completed or terminated trials.
    RESULTS: TrialScout had a sensitivity of 92.5% and a specificity of 81.2% compared to human coders. Manual review of 200 cases where TrialScout disagreed with human researchers showed that a majority (123/200, 61.5%, 95% CI, 54.4-68.3%) of disagreements were due to human errors. When used on 9,600 sampled trials in ClinicalTrials.gov, TrialScout found result publications for 6,110 (63.6%) of trials.
    DISCUSSION: TrialScout reliably located results of completed clinical trials. The tool offers benefits in terms of speed and efficiency. Estimating TrialScout's accuracy is limited by the lack of a true gold standard. TrialScout can accelerate the process of locating trial results in the scientific literature and can assist in monitoring trial reporting practices.
    Keywords:  Clinical Trials; Data Science; Evidence Synthesis; Publication Bias; Publication Rates; Research Transparency
    DOI:  https://doi.org/10.1016/j.jclinepi.2026.112484
  15. Med Teach. 2026 Aug 26. 1-13
       INTRODUCTION: Large language models (LLMs) are increasingly proposed as deductive coders in qualitative research, but their measurement properties remain underexplored. This study applies generalizability theory to evaluate whether hybrid human-LLM workflow configurations can achieve reliable mode-outcomes as an alternative consensus-generating mechanism for deductive coding tasks in medical education research.
    METHODS: Three commercial LLMs (GPT-5.2, Claude Opus 4.5, Gemini 3-Flash Preview) coded 741 excerpts from a published audit of AI-related policy documents at 146 U.S. medical schools against a 24-subtheme deductive framework. Mixed-effects logistic regression assessed variability in agreement with human consensus across coder type (human versus LLM), excerpt characteristics (complexity and length), and coding conditions (sequential independent versus batched processing). A simulation-based D-study forecasted agreement levels for various hybrid human-LLM configurations.
    RESULTS: Sequential independent LLM coding showed agreement comparable to human coding (β = -0.14, p = 0.326). Batched processing showed significantly lower human-LLM agreement (β = -0.41, p = 0.007) and substantial batch-related variance (variance = 30.16). The three LLMs produced similar agreement (joint Wald χ2 = 0.15, p = 0.930), and LLM-LLM agreement (κ = 0.750-0.755) substantially exceeded human-human agreement (κ = 0.422) and human-versus-consensus agreement (κ = 0.420-0.533). LLM disagreements were more systematically patterned across the codebook (Cramer's V = 0.554-0.585) than human disagreements (Cramer's V = 0.196). Simulations forecasted agreement ranging from 0.520 to 0.529 when two or more LLMs were paired with one to two human coders. These forecasted agreement levels for hybrid human-LLM workflows fell within the range of human-versus-consensus agreement.
    DISCUSSION: D-study simulations support the use of hybrid human-LLM workflows to reach coding consensus for deductive reasoning tasks through mode responses, an alternative consensus-generating mechanism to traditional adjudication discussions. Future work should examine whether these patterns extend across additional deductive coding contexts and model families.
    Keywords:  Artificial intelligence; deductive coding; education; medical education; qualitative research
    DOI:  https://doi.org/10.1080/0142159X.2026.2721362
  16. Ther Innov Regul Sci. 2026 Aug 26.
      Artificial intelligence (AI) has the potential to strengthen real-world evidence (RWE) for regulatory decision-making, but its contribution varies by application and methodological maturity. RWE remains limited by challenges in data quality, population selection, treatment characterization, outcome assessment, and statistical methodology. Machine learning and generative AI (genAI), combined with causal inference frameworks, may address these challenges. We review applications, limitations, including reproducibility, bias, transportability, uncertainty quantification, and regulatory acceptability, and methodological priorities for fit-for-purpose AI-enabled RWE.
    Keywords:  AI; Causal inference; Dataset shift; Decision making; RWE; Regulatory; Transportability
    DOI:  https://doi.org/10.1007/s43441-026-01036-5
  17. Med Sci Monit. 2026 Aug 29. 32 e953782
      BACKGROUND Large language models (LLMs) are increasingly used in healthcare; concerns persist regarding the accuracy of generated bibliographic references. The effect of prompt design on reference reliability has not been clearly established. This comparative experimental study evaluated the impact of prompt specificity on LLM-generated reference accuracy in endodontics and compared model performance. MATERIAL AND METHODS We used ChatGPT 5 and Claude Sonnet 4.6. Ten predefined endodontic queries were combined with 3 prompt types of increasing specificity. Each model generated 5 references per query-prompt combination (total: 300 references). References were verified using PubMed, Google Scholar, and CrossRef. Accuracy was classified as fabricated (0), partially accurate (1; existing references containing ≥1 bibliographic inaccuracy), or fully accurate (2). Digital object identifier (DOI) accuracy was assessed separately. Statistical analyses were performed using mixed-effects models and Pearson's chi-square test or Fisher's exact test. RESULTS Accuracy scores tended to increase with greater prompt specificity (P=0.249). Claude demonstrated significantly higher accuracy than ChatGPT (mean score: 1.79 vs 1.25; P<0.001). DOI accuracy did not differ among prompt groups (P=0.338); it was significantly higher for Claude than for ChatGPT (90.0% vs 35.3%; P<0.001). ChatGPT produced significantly more title, journal, and DOI errors (P<0.001); author and year errors were similar between models. CONCLUSIONS Prompt specificity had limited effects on reference accuracy; model selection played a greater role. DOI accuracy was strongly model-dependent and largely unaffected by prompt design under the test conditions, highlighting the need for external verification of LLM-generated references.
    DOI:  https://doi.org/10.12659/MSM.953782
  18. J Med Internet Res. 2026 Aug 26. 28 e96002
       BACKGROUND: Rehabilitation clinical practice guidelines (CPGs) have increased rapidly, but inconsistent methodological quality limits their implementation. Although Appraisal of Guidelines for Research and Evaluation II (AGREE II) and Reporting Items for Practice Guidelines in Health Care (RIGHT) provide standardized appraisal frameworks, their application is time-consuming. Large language model (LLM)-based AI agents may offer a scalable alternative with uncertain reliability.
    OBJECTIVE: We evaluated rehabilitation CPGs' methodological and reporting quality and determined whether structured guidance improves human expert-AI agent agreement.
    METHODS: We systematically reviewed English- and Chinese-language rehabilitation CPGs from Embase, Scopus, PubMed, China National Knowledge Infrastructure, Wanfang Data, National Institute for Health and Care Excellence, Scottish Intercollegiate Guidelines Network, and Guidelines International Network up to June 2026. Methodological and reporting quality were assessed using AGREE II and the RIGHT checklist. Factors associated with guideline quality were examined using regression and subgroup analyses. Two AI agents were compared with human consensus with and without a structured guideline appraisal workbook, followed by external validation using 6 anterior cruciate ligament reconstruction CPGs.
    RESULTS: We included 227 CPGs (163 English-language, 64 Chinese-language). After introducing a structured guideline appraisal workbook, agreement among human experts improved markedly-mean intraclass correlation coefficients (ICCs) increased from -0.09 to 0.66 to 0.84-0.92 across AGREE II domains. Overall guideline quality remained low, with 35.9% (SD 18,8%) applicability and 52% (SD 17.2%) stakeholder involvement. English-language guidelines outperformed Chinese-language guidelines in scope and purpose (mean 74.64, SD 15.4 vs mean 68.88, SD 13.9; P=.004) and applicability (mean 39.14, SD 18.6 vs mean 27.54, SD 16.7; P<.001). Backward-elimination logistic regression revealed external review as an associated process characteristic (odds ratio 20.39, 95% CI 4.66-89.27; P<.001). RIGHT assessments showed consistent reliability (ICC=0.80-0.88). Reporting was highest for basic information (70.5%) and lowest for funding, declaration, and management of interests (44.1%). Meta-analysis of RIGHT reporting rates showed lower reporting among Chinese-language than English-language guidelines (risk difference [RD] -0.07, 95% CI -0.13 to -0.02, 95% prediction interval [PI] -0.39 to 0.24) and among guidelines published before vs after RIGHT release (RD -0.19, 95% CI -0.26 to -0.13, 95% PI -0.56 to 0.17). Without additional guidance, agent-human agreement was moderate (ICC=0.608-0.629). The workbook improved agreement for both models, with DeepSeek-R1's increasing from 0.613 to 0.709 and o1-mini's from 0.629 to 0.687. In validation beyond rehabilitation, DeepSeek-R1 maintained stable agreement (ICC=0.711) and completed appraisals in 5.44 minutes compared to 11.18 minutes for humans.
    CONCLUSIONS: Rehabilitation CPGs, particularly Chinese-language CPGs, continue showing deficiencies in applicability and stakeholder involvement. LLM-based appraisal without structured guidance provides insufficient agreement. Structured guidance improved agent-human agreement, supporting AI-assisted guideline appraisal under human oversight. Although further validation across additional clinical specialties is needed, AI agents can serve as efficient assistants in guideline appraisal instead of replacing humans. Future synthesis requires human-AI integration guided by structured, expert-defined principles.
    TRIAL REGISTRATION: PROSPERO CRD420251270676; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251270676.
    Keywords:  AGREE II; RIGHT; large language models; practice guideline; rehabilitation
    DOI:  https://doi.org/10.2196/96002
  19. NPJ Digit Med. 2026 Aug 27. pii: 662. [Epub ahead of print]9(1):
      In this Reply, we respond to the Matters Arising by Zhang and Fu regarding the trustworthiness of AI-assisted clinical guideline development, using Quicker as a case study. We clarify which safeguards for transparency, traceability, and uncertainty handling are already embedded, to a substantial extent, in Quicker's design, and outline areas of alignment on validation, governance, and modular evaluation. We further discuss broader challenges related to trust, safeguards, and responsible deployment of large language model-based systems in evidence-based medicine. We emphasize that advancing such systems requires not only technical acceleration, but also systematic human oversight, rigorous evaluation frameworks, and community-wide governance to ensure safe and trustworthy clinical adoption.
    DOI:  https://doi.org/10.1038/s41746-026-03099-y
  20. Clinicoecon Outcomes Res. 2026 ;18 594790
      An infographic titled 'What to Expect in 2026-2027: Important Health Economics and Outcomes Research (HEOR) Topics' outlines five key HEOR areas. 1. HEOR Models: AI is transforming HEOR with advanced data, but needs validation to avoid errors. 2. Real World Evidence: Essential in healthcare, it connects clinical trials with real-life practice for improved access and regulation. 3. Prevention Economics: Now a policy focus, showing that investing in prevention is financially crucial for effective implementation. 4. Health Equity: Involves expanding access, integrating social determinants and developing methods to measure healthcare impact. 5. Healthcare Resource Reallocation: Spending is shifting from austerity to resilience, emphasizing prevention for better fiscal outcomes than defense spending. The infographic includes decorative elements.An infographic listing top five HEOR topics for 2026-2027, numbered 1 to 5 with headings and summaries.
    DOI:  https://doi.org/10.2147/CEOR.S594790
  21. J Endod. 2026 Aug 28. pii: S0099-2399(26)00464-4. [Epub ahead of print]
       BACKGROUND: Large language models (LLMs) are increasingly used by clinicians and learners for endodontic information, yet their reliability for irrigation-related knowledge remains unclear. This study evaluated the accuracy, readability, hallucination profile, clinical risk, and temporal stability of 4 general-purpose LLMs on endodontic irrigation questions, including false-premise prompts.
    METHODS: A literature-based reference set was developed for sodium hypochlorite, calcium hypochlorite, ethylenediaminetetraacetic acid, and chlorhexidine. ChatGPT-5.2, Claude Sonnet 4.5, Gemini 3 Pro, and DeepSeek 3.2 were each asked 100 questions comprising 80 factual items and 20 contradiction-seeking items. Responses were independently scored by 2 blinded endodontists for accuracy, hallucination subtype, and clinical risk. Five readability indices were calculated, and model stability was reassessed 10 days later.
    RESULTS: Inter-rater agreement was almost perfect (weighted κ = 0.92 for accuracy; κ = 0.97 for hallucination). Claude Sonnet 4.5 and Gemini 3 Pro showed the highest accuracy (1.75 and 1.74) and the lowest hallucination rates (7% and 9%). DeepSeek 3.2 showed the lowest accuracy (1.08), the highest hallucination rate (29%), and critical-risk outputs in 16% of responses. Gemini showed the highest test-retest stability (weighted κ = 0.95). Hallucination strongly correlated with clinical risk (ρ = 0.93; P < 0.001). Readability analysis showed a 2-tier pattern: Gemini and DeepSeek produced more accessible text, whereas ChatGPT and Claude generated denser outputs.
    CONCLUSIONS: LLM performance in endodontic irrigation is model- and irrigant-dependent. Hallucination profiling is clinically relevant, and high initial accuracy does not guarantee temporal stability. LLM outputs should be used only as clinician-verified adjuncts.
    Keywords:  Artificial intelligence; calcium hypochlorite; endodontic irrigation; hallucination; large language models; sodium hypochlorite
    DOI:  https://doi.org/10.1016/j.joen.2026.08.033