bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–07–26
eighteen papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. Campbell Syst Rev. 2026 Sep;22(3): 18911803261469813
       Background: Natural language processing (NLP) techniques offer promising solutions for semi-automating the time-consuming process of abstract screening in systematic reviews. The exponential growth of published literature has created significant bottlenecks, with review teams manually assessing thousands of abstracts over weeks to months. Single reviewers can miss 5-13% of relevant studies, necessitating dual screening that further increases workload. Advances in artificial intelligence, including deep learning models such as BERT and its successors, show potential for automating this critical step, but comprehensive evidence on optimal approaches, performance, and practical feasibility remains limited.
    Objectives: This systematic review aimed to assess techniques, performance, and feasibility of NLP approaches for title and abstract screening by characterizing the range of NLP methods used, summarizing performance on key metrics like workload reduction and recall, evaluating real-world implementation feasibility, and identifying research gaps and future directions.
    Search Methods: We searched PubMed, Web of Science, Embase, CINAHL, The Cochrane Library, Scopus, and gray literature sources from inception to December 2024. The search strategy, developed with an information specialist and peer-reviewed using PRESS guidelines, targeted keywords related to natural language processing, machine learning, abstract screening, and systematic reviews. Additional sources included conference proceedings, preprint servers, reference lists, forward citation tracking, and expert consultation.
    Selection Criteria: We included primary studies of any design describing development or evaluation of NLP techniques for automating title and abstract screening in evidence syntheses. Eligible studies reported on NLP methods, screening performance (workload reduction, recall, precision), or implementation feasibility. Studies using only rule-based approaches without machine learning, systematic reviews of NLP methods, commentaries, and conference abstracts were excluded. No language or date restrictions were applied.
    Data Collection and Analysis: Two reviewers independently screened titles, abstracts, and full texts using Covidence software, with disagreements resolved through discussion. Data extraction covered study characteristics, NLP techniques, training approaches, performance metrics, and feasibility considerations. Risk of bias was assessed using a modified ROBIS tool. Given diverse techniques and outcomes, we conducted narrative synthesis following SWiM guidelines, grouping studies by NLP approach.
    Main Results: From 4,105 records, 19 studies met inclusion criteria, with 68.4% published since 2023, reflecting rapid field advancement. Studies employed diverse approaches from traditional machine learning (Support Vector Machines, Random Forests) to advanced deep learning models, particularly BERT variants. Most achieved >90% recall with workload reductions of 13-96%, representing substantial time savings. Deep learning models with transfer learning consistently outperformed traditional approaches. However, implementation faced significant barriers including requirements for high-quality training data, specialized computational resources, technical expertise, and user-friendly interfaces. Performance was generally better for targeted reviews with lower inclusion prevalence.
    Authors’ Conclusions: NLP techniques, especially deep learning with transfer learning, show substantial promise for semi-automating abstract screening with potential for large workload savings while maintaining high recall. However, challenges remain regarding training data quality, computational requirements, technical expertise needs, and user-centered design. Realizing full potential requires interdisciplinary collaboration to develop reliable, generalizable tools integrating seamlessly with human expertise and existing workflows. Future priorities include creating standardized datasets, conducting prospective evaluations, developing user-friendly interfaces, and establishing implementation best practices to revolutionize evidence synthesis efficiency.
    Keywords:  abstract screening; deep learning; evidence synthesis; machine learning; natural language processing; systematic review automation
    DOI:  https://doi.org/10.1177/18911803261469813
  2. Cureus. 2026 Jun;18(6): e111193
      Background The growth of biomedical literature increasingly exceeds the capacity of manual evidence synthesis. Large language models (LLMs) may support abstract screening and structured extraction, but many current workflows depend on proprietary cloud APIs, creating challenges for governance, reproducibility, and scalable deployment. Methods I developed a fully local, open-weight, schema-constrained pipeline (gpt-oss-20b, deployed via Ollama on Apple M1 Max) for title/abstract-based scoping workflows. The pipeline combined deterministic metadata filtering, LLM-assisted screening, and structured abstract extraction. Performance was benchmarked against three published systematic reviews (ketamine/neuroimaging; clozapine/suicidality; clozapine patient/caregiver perspectives) using precision, recall, and F1 against reference inclusion sets. I also report audit-adjusted estimates (i.e., performance metrics recalculated after manual full-text adjudication of discrepant records) alongside standard reference-set performance. Results In the ketamine/neuroimaging benchmark, the pipeline retained all 41 studies included in the original review; after audit adjustment, recall was 100.0% (46/46), accuracy 99.4% (156/157), precision 97.9% (46/47), and F1 98.9%. For clozapine/suicidality, recall was 79.3% (46/58), and F1 was 76.0%, with missed studies largely attributable to missing or non-informative abstracts. For clozapine patient/caregiver perspectives, recall was 88.9% (56/63), and F1 was 83.6%, with similar abstract-level constraints. Abstract-level extraction recovered audited metadata fields without detected errors and generated evidence maps that were thematically concordant with the main narrative structure of the reference reviews. Conclusions As a proof-of-concept, a fully local LLM pipeline can support scalable and auditable abstract-based scoping and high-level evidence mapping. Because performance was benchmarked against three reviews with partly audit-adjusted reference sets, the findings require confirmation in larger, independently adjudicated evaluations. Random human audit remains advisable, and expert full-text synthesis remains necessary when abstracts are non-informative or when mechanistic precision is required.
    Keywords:  brain imaging; large language models; mapping evidence; psychiatry; scoping review
    DOI:  https://doi.org/10.7759/cureus.111193
  3. J Evid Based Med. 2026 Jul 25. e70166
       OBJECTIVE: To systematically evaluate the diagnostic performance of large language models (LLMs) in automated medical literature screening and to determine their potential role in supporting evidence synthesis workflows.
    METHODS: PubMed, Web of Science, Embase, The Cochrane Library, Google Scholar, CNKI, Wanfang, VIP, and CBM were searched from January 1, 2022 to June 11, 2026. Studies assessing LLMs for automated title and abstract screening or full-text eligibility assessment in medical literature were included. The primary outcomes were sensitivity and specificity. Secondary outcomes included positive and negative likelihood ratios, diagnostic odds ratio, area under the curve (AUC), and efficiency-related metrics. Pooled sensitivity and specificity were estimated using a bivariate random-effects model and hierarchical summary receiver operating characteristic framework. Subgroup analyses and meta-regression were performed to explore sources of heterogeneity. This systematic review was registered with the Open Science Framework (https://osf.io/56d8j).
    RESULTS: Eighteen studies published between 2023 and 2025 were included. In title and abstract screening, the pooled sensitivity was 0.92 (95% confidence interval [CI]: 0.81-0.96) and pooled specificity was 0.94 (95% CI: 0.90-0.97). The summary receiver operating characteristic AUC reached 0.98 (95% CI: 0.96-0.99). In full-text screening, pooled sensitivity and specificity both reached 0.99 (95% CI: 0.95-1.00) and the AUC was 0.99 (95% CI: 0.98-1.00). Prompt strategies incorporating examples or chain-of-thought reasoning were associated with higher sensitivity than strategies without these approaches (0.95 vs. 0.86, p < 0.01). Several studies reported substantial efficiency gains, including workload reductions ranging from approximately 50% to 99% and screening time reductions of up to tenfold.
    CONCLUSIONS: LLMs shows promising performance in automated medical literature screening, particularly in full-text assessment. These models show strong potential as high sensitivity assistive tools that can substantially reduce manual screening burden while supporting evidence synthesis. Further real-world validation is needed to establish their role in evidence-based medicine.
    Keywords:  large language model; literature screening; meta‐analysis
    DOI:  https://doi.org/10.1111/jebm.70166
  4. NAM J. 2026 ;2 100112
      The validation of in vitro New Approach Methodologies (NAMs) requires the use of well-characterized reference chemicals to assess assay performance, reproducibility, and relevance. However, selecting such chemicals is often labor-intensive and lacks standardization. We developed a semi-automated, evidence-based workflow that integrates systematic literature review, AI-assisted data extraction, and quantitative evidence scoring to improve the efficiency and transparency of chemical selection. We conducted a structured EMBASE search, followed by AI-assisted abstract screening and data extraction, in a case study to select reference chemicals for the validation of an in vitro adipogenesis assay. Chemicals were evaluated for in vitro, in vivo, and human evidence of adipogenic or obesogenic effects. From 11,648 screened publications, 236 studies met the inclusion criteria, resulting in the identification of 243 candidate reference chemicals, of which 50 were prioritized based on scoring. The final selection encompassed a set of 22 chemicals with different levels of potency and diverse structures. This workflow illustrates how AI-assisted evidence synthesis can accelerate and standardize the selection of reference chemicals while maintaining expert oversight and ensuring regulatory relevance. It provides a reproducible and adaptable framework for future NAMs validation studies across diverse toxicological endpoints.
    Keywords:  Adipogenesis; Artificial intelligence; Assay validation; Chemical selection; Large language models (LLMs); New approach methodologies (NAMs); Reference chemicals
    DOI:  https://doi.org/10.1016/j.namjnl.2026.100112
  5. Open Forum Infect Dis. 2026 Jul;13(7): ofag401
       Background: Rapid evidence synthesis during emerging infectious and re-emerging disease outbreaks is critical, yet traditional systematic reviews rarely meet urgent timelines. Large language models (LLMs) may accelerate evidence synthesis by extracting data from publications. We compared an LLM-assisted data extraction system with manual extraction.
    Methods: We conducted a 1:1, open-label, 2-period, randomized crossover trial at the National Center for Global Health and Medicine, a national reference center for emerging infectious diseases in Japan (2025). Five experienced reviewers extracted predefined items from mpox-related articles under 2 conditions: (i) LLM-assisted extraction using OpenAI's o3 model to generate structured summaries and (ii) manual review of PDF files. The primary outcome was task completion time; secondary outcomes were extraction accuracy and adverse events. Mixed-effects models included condition as a fixed effect and participant and paper IDs as random effects. The protocol, source code, and data are available at https://github.com/SRWS-PSG/emerging_infection_24K13518_open.
    Results: Five evaluators (4 physicians and 1 pharmacist; 6-10 years postgraduation) completed 20 task-level evaluations (LLM, n = 9; no LLM, n = 11). Mean completion time was 27.5 minutes with LLM assistance versus 34.5 minutes without. The LLM-assisted condition was 7.9 minutes faster on average (95% CI -1.5 to 17.3; P = .099). Extraction accuracy was 100% in both conditions, and no adverse events were reported.
    Conclusions: LLM assistance might reduce data extraction time by ∼23% (7.9 minutes per article; 95% CI -1.5 to 17.3 minutes) with no observed loss of accuracy. Although statistical uncertainty remains, LLM integration may offer practical value for rapid evidence synthesis during public health emergencies as tools and prompting strategies mature.
    Keywords:  artificial intelligence; emerging infections; large language model; randomized crossover trial
    DOI:  https://doi.org/10.1093/ofid/ofag401
  6. Campbell Syst Rev. 2026 Sep;22(3): 18911803261469865
      
    Keywords:  Artificial intelligence; Evidence synthesis; Information specialists; Large language models; Prompt engineering; Search strategy
    DOI:  https://doi.org/10.1177/18911803261469865
  7. JBI Evid Synth. 2026 Jul 22.
       OBJECTIVE: The objective of this study was to examine how sensitivity and specificity of title/abstract screening decisions change as the proportion of records that are co-screened is varied.
    INTRODUCTION: With an accelerating volume of literature being published, the demand for systematic reviews is increasing. To maintain the feasibility of conducting reviews in a timely fashion, efficiencies must be found in the review process that do not compromise their rigor. Current guidelines recommend that all records be dual-screened. Less is known about the impact of co-screening a subset of records on screening accuracy.
    METHODS: Eight reviewers independently screened the titles and abstracts from a published systematic review and meta-analysis. Using their decisions, simulations were run to generate permutations of reviewer combinations in 2 formats: dual screening (ie, primary reviewer screened all records, secondary reviewer screened a portion, third reviewer mediated disagreements) and 3-way screening (ie, primary reviewer, secondary reviewer, and mediator roles were equally divided between 3 individuals). Within each permutation, further simulations were created so that the secondary reviewer screened increasing proportions of records in 5% increments. Decisions were compared against the final full-text decisions used in the published review to determine sensitivity and specificity of decisions for each permutation, at each increment. The relationship between sensitivity/specificity and percentage of co-screening was assessed using linear mixed-effects models.
    RESULTS: All reviewers screened 3891 potentially eligible records, of which 80 were included in the final review. Sensitivity (median range: 93.3%-96.3%) and specificity (median range: 87.7%-89.6%) were stable across different proportions of co-screening, for both dual- and 3-way screening strategies. Sensitivity was higher in reviewer groupings where at least one reviewer had subject matter experience. Specificity tended to be higher when the primary reviewer had experience in conducting a systematic review.
    DISCUSSION: Our results suggest that complete co-screening of title/abstract records does not appreciably improve sensitivity or specificity. Rather, even low levels of co-screening across dual- and 3-way screening can produce high accuracy. Reviewer experience in subject matter and in conducting a systematic review, however, appear to be influential. These results are based on simulations and permutations from a single review so should be used as a guide only.
    CONCLUSION: Co-screening a proportion of title/abstracts in a systematic review appears to produce comparable sensitivity and specificity to co-screening all records. Proportional co-screening may therefore represent an opportunity for enhancing efficiency in the systematic review process.
    REVIEW REGISTRATION: https://osf.io/9feuh.
    Keywords:  record screening; study within a review; systematic review
    DOI:  https://doi.org/10.11124/JBIES-25-00309
  8. Value Health. 2026 Jul 18. pii: S1098-3015(26)02557-X. [Epub ahead of print]
       OBJECTIVES: Large language models (LLMs) may be useful tools for the development/adaptation of patient-reported outcome measures (PROMs). As a methodological proof-of-concept, we evaluated the use of LLMs to support the identification of potential EQ-5D-5L bolt-on dimensions, using patient-reported free-text data.
    METHODS: We used GPT-4o to analyze text data from 1,977 members of the Dutch Celiac Association, who completed the EQ-5D-5L and narratively described the impact of celiac disease on their lives. Prompts were designed to identify potential EQ-5D-5L bolt-on dimensions and produce preliminary item wordings for selected dimensions. Evaluations comprised: comparisons of dimensions identified using two alternative approaches (qualitative analysis and topic modelling) conducted on a subset of 85 text entries; text-entry level agreement (Kappa) between LLM and qualitatively identified dimensions; and suitability of LLM-generated item wordings assessed against existing criteria.
    RESULTS: The LLM identified 12 potential bolt-on dimensions to the EQ-5D-5L, of which 9 were also identified using qualitative analysis, and 5 using topic modelling. Text-entry level agreement between the LLM and qualitative approaches was 'moderate', 'substantial' or 'almost perfect', with two exceptions of slight/fair agreement (median Kappa=0.68, IQR=0.56-0.71). Sensitivity analyses using four other LLMs produced similar agreement results. The LLM-generated item wordings for the 4 most common dimensions scored 4.0-4.4 out of 5 when assessed against existing criteria.
    CONCLUSIONS: This study demonstrates the potential of LLMs to support the development/modification of PROMs based on patient-reported text data. Further research should assess the approach's transferability across disease areas and data sources, while better incorporating patient/stakeholder input throughout.
    Keywords:  artificial intelligence; large language models; patient reported outcomes; preference based measures
    DOI:  https://doi.org/10.1016/j.jval.2026.06.021
  9. Nat Protoc. 2026 Jul 24.
      Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4 and DeepSeek-R1, represent a transformative class of artificial intelligence tools capable of revolutionizing various aspects of healthcare by generating human-like responses across diverse contexts and adapting to novel tasks following human instructions. Their potential application spans a broad range of medical tasks, such as clinical documentation, matching patients to clinical trials and answering medical questions. Here in this Tutorial, we discuss an actionable set of best practices to help healthcare professionals utilize LLMs more effectively and efficiently. The overall workflow follows sequential phases from formulating the task, choosing the most appropriate LLMs, engineering the prompts, fine-tuning the requests and through to model deployment. We discuss a set of critical considerations in identifying medical tasks that align with the core capabilities of LLMs and selecting models based on the required task, data, performance and model interface. We then review the strategies, such as prompt engineering and fine-tuning, to adapt standard LLMs to specialized medical tasks. We then cover deployment considerations, including regulatory compliance, ethical guidelines and continuous monitoring for fairness and bias. By providing a structured step-by-step methodology, this entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
    DOI:  https://doi.org/10.1038/s41596-026-01408-z
  10. Front Res Metr Anal. 2026 ;11 1833223
      Fraudulent, falsified, and otherwise unreliable clinical research threatens the foundations of evidence-based medicine because systematic reviews, meta-analyses, and clinical practice guidelines depend on the trustworthiness of the studies they include. Retractions are increasing faster than publication output and often occur only after substantial delay, allowing problematic trials to be cited, pooled, and incorporated into downstream clinical recommendations. Recent evidence shows that retracted randomized trials have contaminated thousands of meta-analyses and hundreds of guideline documents, while only a small minority of affected reviews later self-correct. This is not only a clinical problem but also an integrity analytics problem: unreliable studies generate article-level, trial-level, and network-level signals that are not yet systematically integrated into evidence-synthesis workflows. Existing approaches can detect some image anomalies, textual irregularities, reporting inconsistencies, and statistical red flags, but important blind spots remain. Fabricated clinical datasets may appear statistically plausible, outcome switching often requires registration-publication comparison, and most available tools operate in isolation rather than as coordinated screening systems. We argue that the next generation of safeguards should combine AI/ML-based detection systems, statistical forensic methods, structured trustworthiness appraisal tools, and workflow-level screening frameworks within evidence synthesis. Embedding staged trustworthiness assessment early in systematic reviews, alongside stronger publisher- and database-level integrity infrastructure, could help prevent unreliable trials from distorting pooled estimates and downstream guidance. Such systems should support triage and prioritization rather than replace human judgment. Framed in this way, scalable trustworthiness assessment represents a practical research metrics and analytics agenda for strengthening the evidence ecosystem from primary studies to reviews, guidelines, and policy decisions.
    Keywords:  artificial intelligence; evidence synthesis; integrity analytics; machine learning; research integrity; retractions; systematic reviews; trustworthiness assessment
    DOI:  https://doi.org/10.3389/frma.2026.1833223
  11. Regul Toxicol Pharmacol. 2026 Jul 24. pii: S0273-2300(26)00155-8. [Epub ahead of print] 106182
      Risk assessment activities at the European Food Safety Authority face mounting challenges from an increasing volume of compounds requiring evaluation and exponential growth in scientific literature. To address these challenges, toxicologists started grouping compounds based on their mechanism of action in the body, allowing study comparisons similar to current risk assessment strategies. The LLMs-rev pipeline was created to retrieve potentially relevant papers from PubMed, assess their relevance, and extract the mechanism of action (MoA) information from accessible publications using large language models. Applied to 121 compounds, the system processed over 400,000 papers, identifying 30,250 as containing relevant information and extracting specific MoA quotations from 4,500 open-access publications. The automated approach demonstrated processing speeds exceeding 8,000 papers per hour, dramatically outpacing conventional manual screening methods that typically assess 100-120 papers per hour. While the system proved particularly effective for compounds for which abundant literature was available, human expertise remained essential for interpreting complex MoAs that required contextual data analysis. The optimized prompts minimized model hallucinations by restricting outputs to direct quotations from verified sources. The methodology demonstrates considerable potential for accelerating risk assessment workflows while maintaining scientific rigor through the complementary use of human oversight.
    DOI:  https://doi.org/10.1016/j.yrtph.2026.106182
  12. World J Otorhinolaryngol Head Neck Surg. 2026 Jun 26.
       Objective: To compare the quality of scientific review articles generated by two artificial intelligence systems, ChatGPT and Gemini, with those written by human authors in the field of otolaryngology.
    Methods: Two otolaryngology topics, chronic rhinosinusitis and infantile subglottic hemangioma, were selected. For each topic, four AI-generated reviews (GPT-4.0 and Gemini 2.0; narrative and PRISMA-style) and one human-authored peer-reviewed review were included, yielding a total of 10 manuscripts (8 AI-generated, 2 human-authored). A blinded panel of seven board-certified otolaryngologists evaluated all manuscripts using a 5-point Likert scale across seven domains: scientific accuracy, depth of content, citation quality, structure and organization, readability and tone, critical insight, and overall scientific quality. Group comparisons were performed using linear mixed-effects models with random intercepts for reviewer and manuscript. Interrater reliability was assessed using Shrout-Fleiss intraclass correlation coefficients (ICC). Manual verification of AI-generated references was conducted to assess citation accuracy and fabrication.
    Results: Human-authored manuscripts received the highest ratings across all domains (overall quality 4.50 ± 0.76). GPT-4.0 demonstrated moderate performance (2.71 ± 1.46 overall), while Gemini 2.0 scored lowest (2.14 ± 1.01). Mixed-effects modeling demonstrated significant group differences across all domains (p ≤ 0.008). Citation quality showed one of the largest between-group differences and strong reliability [ICC(2,1) = 0.68; ICC(2,7) = 0.94]. Manual verification of 123 AI-generated references revealed high citation accuracy for GPT-4.0 (89.6% fully accurate; 0% fabricated) compared with Gemini 2.0 (71.7% fully accurate; 23.9% fabricated). Reviewers misclassified 50% of GPT-4.0 manuscripts as human-authored, correctly identified 93% of human-authored manuscripts, and classified 86% of Gemini manuscripts as AI-generated.
    Conclusion: GPT generated fluent, stylistically strong reviews but remained significantly inferior to human-authored manuscripts in analytical depth and citation integrity. Gemini 2.0 underperformed across all domains and demonstrated a substantial rate of fabricated citations. As large language models become integrated into academic workflows, transparent disclosure, structured fact-checking, and human oversight remain essential to safeguard scientific reliability.
    Keywords:  GPT‐4.0; citation integrity; generative artificial intelligence; otolaryngology; peer review; scientific writing
    DOI:  https://doi.org/10.1002/wjo2.70131
  13. JBI Evid Implement. 2026 Jul 23.
      Implementation science aims to bridge the gap between research evidence and routine health care practice by understanding and optimizing the integration of evidence-based interventions. In this paper, we identify seven persistent challenges limiting implementation progress, including (1) overwhelming volume of implementation materials (e.g., reports, interviews, surveys); (2) contextual variability; (3) complex interactions between contextual factors, interventions, and outcomes; (4) interest holder engagement constraints; (5) equity and access barriers; (6) insufficient or biased data; and (7) concerns around data security, trust, and ethics. We explore how advances in data science and artificial intelligence (AI) offer promising solutions to these challenges by enhancing evidence extraction and synthesis, contextual analysis, interest holder engagement, and adaptation throughout the implementation process. Using a live project in precision oncology as a practical example, we demonstrate how AI tools, such as large language models, clustering algorithms, and sentiment analysis, can support implementation research and practice, from concept (e.g., barriers, facilitators, strategies), extraction, and barrier identification to process mapping. We also introduce ImpleMATE, an AI-enabled platform that integrates implementation science knowledge with dynamic learning health system workflows, enabling continuous knowledge extraction, decision support, and feedback for implementation improvement. While AI offers significant potential to accelerate and scale up implementation efforts, we emphasize the need for ethical oversight, transparency, and human collaboration to ensure responsible, equitable, and impactful application in practice.
    SPANISH ABSTRACT: http://links.lww.com/IJEBH/A661.
    Keywords:  artificial intelligence; data science; evidence-based intervention; implementation science; large language model
    DOI:  https://doi.org/10.1097/XEB.0000000000000635
  14. J Clin Epidemiol. 2026 Jul 18. pii: S0895-4356(26)00300-8. [Epub ahead of print] 112424
      
    Keywords:  CHA studies; LLMs; chatbots; generative AI; large language models; regulation
    DOI:  https://doi.org/10.1016/j.jclinepi.2026.112424
  15. JACC Adv. 2026 Jul;pii: S2772-963X(26)00339-X. [Epub ahead of print]5(7): 102918
       BACKGROUND: Cardiac amyloidosis (CA) is increasingly recognized in clinical practice. Whether a guideline-based large language model can deliver clinician-level answer quality for CA remains unknown.
    OBJECTIVES: This study aimed to develop a guideline-based custom generative pretrained transformer (GPT) (AmyloGPT) and evaluate whether its response quality matches or exceeds that of board-certified cardiologists for questions regarding CA.
    METHODS: AmyloGPT was built in OpenAI's GPT Builder without programming, integrating the 2020 Japanese Circulation Society CA guidelines as its knowledge base. Ten nonspecialist physicians generated 71 unique clinical questions. Five board-certified cardiologist answerers drafted responses. In a prospective, blinded, comparative study, evaluators (10 nonspecialists and 3 board-certified cardiologists) assessed paired responses for preference (forced-choice) and response quality using five-point Likert scales.
    RESULTS: Compared with cardiologist answers, AmyloGPT was preferred in 81.1% (95% CI: 78.1%-83.8%) of evaluations by nonspecialist evaluators and 83.6% (95% CI: 78.6%-88.6%) of those by cardiologist evaluators (both P < 0.001). Among nonspecialists, AmyloGPT received higher median ratings for intent alignment and clinical usefulness (both P < 0.001). Among cardiologist evaluators, AmyloGPT received higher median ratings across all 5 quality dimensions: accuracy, consistency, validity, completeness, and absence of bias (all P < 0.001).
    CONCLUSIONS: A no-code, guideline-based custom GPT delivered superior response quality to that of cardiologists for CA questions. This approach allows clinicians without programming skills to build disease-specific large language models, potentially supporting equitable care where specialist access is limited. However, further studies are needed to evaluate potentially inaccurate outputs such as hallucinations.
    Keywords:  GPT-4; artificial intelligence; cardiac amyloidosis; clinical decision support; large language model; practice guidelines
    DOI:  https://doi.org/10.1016/j.jacadv.2026.102918
  16. JMIR Med Inform. 2026 Jul 24. 14 e87831
       BACKGROUND: Medical information extraction requires automatically identifying disease names and related terms in text. This task, known as named entity recognition (NER), relies on expert-annotated data that are costly to produce and often available only in limited quantities. Data augmentation (DA) aims to expand available training data; however, standard techniques such as synonym replacement and back-translation may introduce inappropriate substitutions or fail to preserve entity-label alignment, which is critical for sequence-labeling tasks. Although large language models can generate fluent text, their outputs may also contain factual inconsistencies or unintended changes if not carefully controlled.
    OBJECTIVE: This study investigated whether persona-driven, document-level DA using a large language model could improve biomedical disease NER performance by generating diverse rephrasings of medical documents while preserving annotated entities.
    METHODS: We designed a DA framework using multiple personas that varied in medical expertise, personality, tone, and narrative style. Using prompting constrained by XML tags, each persona rephrased training documents while aiming to preserve annotated entity spans. We evaluated the framework on 2 biomedical disease NER datasets with complementary roles: RareDis, a low-resource rare disease corpus, and National Center for Biotechnology Information (NCBI) disease, a more general disease benchmark. Semantic fidelity and lexical diversity were measured using BERTScore and Bilingual Evaluation Understudy (BLEU-4), respectively, and personas were grouped into high-, balanced-, and low-fidelity subsets. Biomedical pretrained BioBERT models were fine-tuned and evaluated under multiple settings, including gold-standard (GS) data only, synonym replacement, single-persona augmentation, curated persona subsets, and all-persona augmentation. Performance was assessed using microaveraged entity-level precision, recall, and F1-score, and results were examined at both the overall and individual entity-type levels. Performance values are reported as mean (SD).
    RESULTS: Persona-driven DA improved NER performance over GS-only training in both datasets, with the strongest gains obtained by combining multiple persona-generated variants with GS data. In RareDis, the best result was achieved by the low-fidelity subset (mean F1-score 73.35, SD 0.19 vs baseline 71.22, SD 0.45), while in NCBI disease, the all-personas setting performed best (mean F1-score 89.32, SD 0.26 vs baseline 87.82, SD 0.18). In low-resource experiments, the all-personas and high-fidelity persona settings in NCBI disease exceeded the performance of the model trained on 100% GS data using only 60% of the training data, whereas gains in RareDis were more modest. Entity-level analysis showed improvements across RareDis categories, particularly for symptom, and confusion analysis indicated reduced symptom-sign confusion under augmentation.
    CONCLUSIONS: Persona-driven DA improved biomedical disease NER by introducing controlled linguistic variation while largely preserving annotated entities. The strongest gains were obtained when multiple persona-generated variants were combined with GS data, although the benefit varied across datasets. These findings suggest that this approach is a promising strategy for low-resource biomedical NER.
    Keywords:  LLM; NER; NLP; data augmentation; disease; large language model; low-resource; named entity recognition; natural language processing; persona; rare disease
    DOI:  https://doi.org/10.2196/87831
  17. PLoS One. 2026 ;21(7): e0354358
      Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients and other stakeholders together. By understanding AI as 'sociotechnical' where the social and the technical nature of the work and the models are inseparable, we explore the AI development workflow and how stakeholders navigate the challenges and tensions of sharing and generating knowledge across disciplines. We conducted an inductive thematic analysis of 13 semi-structured interviews with participants in early stages of AI-in-healthcare research consortia in the UK. Our findings identify that participants needed to adapt both the tools used for sharing and the information communicated according to their audience, particularly when working with those with a clinical or patient perspective. We identify the novelty of participating in AI research, how AI knowledge is shared, and the inclusion of clinician and patient stakeholder perspectives as key areas within collaborative AI practices in healthcare. These findings highlight that bringing AI into the mix can introduce new obstacles to interdisciplinary work.
    DOI:  https://doi.org/10.1371/journal.pone.0354358