bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–08–30
thirty-one papers selected by
Thomas Krichel, Open Library Society



  1. Med Ref Serv Q. 2026 Aug 27. 1-9
      Time is precious. The ideal goal is to be flexible and have the most efficient and effective methods in organizing all one needs to accomplish. As librarians in a service industry, we do not always own all our time. A phone call, an incoming search request, complicated reference question, or a curbside consult are just some examples that can occupy the time that was supposedly devoted to something else. This column is dedicated to using our precious time wisely. For example: to schedule one's day, delegating, appreciating the value of one's tasks and making high-level decisions. Resources include: Assessment of Time Management Skills (ATMS) a validated awareness assessment for using time to plan and manage daily life tasks, "ritual tasks" or mannerisms to engage before beginning to organize one's time, concepts in time management, including Pomodoro, Eisenhower matrix, the 80/20 rule and "eat that frog." Time management apps such as To Do List, Trello, Notion, ATracker and Toggle Track. Consider this column a boost for one's productivity.
    Keywords:  80/20 rule; Assessment of Time Management Skills (ATMS); Eisenhower matrix; Pomodoro; eat the frog; time management
    DOI:  https://doi.org/10.1080/02763869.2026.2719910
  2. Res Synth Methods. 2026 Aug 27. 1-18
      The standard approach for identifying studies for reviews has remained consistent for many years. This approach can be time-consuming, with quality reliant on searcher skill. Alternate approaches aiming to improve efficiency have been proposed, but there is uncertainty around their performance. To inform searching conduct, we systematically reviewed studies comparing alternate and standard search approaches. We searched PubMed and the Library, Information Science and Technology Abstracts (LISTA) databases on 22 January 2025, browsed the Summarized Research in Information Retrieval for HTA (SuRe Info) website, reviewed reference lists of reviews and book chapters, and conducted a backwards and forwards citation search of included studies. Risk of bias (RoB) was assessed using a modified version of Quality Assessment of Diagnostic Accuracy Studies tool version 2 (QUADAS-2), and alternate approaches were categorised according to whether they potentially (1) improved search conduct, (2) worsened search conduct, or (3) had an impact that was difficult to interpret. We identified 3,520 records and included 14 studies evaluating 12 alternative search approaches. We identified four approaches that potentially improve searches, six which may make searches worse, and two with difficult to interpret results. Most studies had high or unclear RoB arising from the standard search conduct and concern regarding the applicability of the results. Searches may be improved by (1) using seed studies to design searches; (2) adding study design filters; (3) adding backwards citation searching; (4) using automation tools for translating searches across databases. More robust evaluations are needed before we can make stronger recommendations.
    Keywords:  evidence synthesis; information retrieval; methodology; searching; systematic reviews
    DOI:  https://doi.org/10.1017/rsm.2026.10112
  3. Eur Heart J Digit Health. 2026 Aug;7(7): ztag125
       Aims: The rapid expansion of biomedical literature challenges clinicians' and researchers' ability to identify clinically meaningful evidence. We systematically compared five literature search tools, four artificial intelligence (AI)-assisted and one conventional, across clinically relevant cardiology research scenarios, using a blinded expert-validated gold standard to assess their ability to retrieve relevant and key references.
    Methods and results: We evaluated ChatGPT-5, Elicit, Consensus, Scite, and PubMed across four cardiology topics defined by maturity and specificity, with multiple standardized prompts. Three electrophysiology experts independently and blindly rated all retrieved references, defining two gold standards: expert-rated relevance and expert-selected key references. ChatGPT-5 achieved the highest proportion of relevant articles (90% [88-100], P < 0.001) and the highest key-reference overlap (60% [43-68], P < 0.001), whereas Scite performed lowest (20% and 10%, respectively). The tool was the primary determinant of performance (partial R 2 = 0.50), whereas prompt formulation had no significant effect. In a pre-specified subanalysis restricted to clinical studies, ChatGPT-5 and human-conducted systematic reviews overlapped by 42% (96% of shared articles highly relevant), with 58% distinct references, indicating complementary AI and human retrieval; ChatGPT-5 produced hallucinated citations when long reference lists were requested for emerging topics, underscoring the need for human verification.
    Conclusion: AI-assisted tools showed heterogeneous performance, ChatGPT-5 performing best in this cardiology setting. These preliminary, context-specific findings support hybrid human-AI strategies in which AI complements rather than replaces transparent database searches such as PubMed; larger-scale, multi-domain studies are needed to confirm and generalize them.
    Keywords:  Artificial intelligence; Cardiac resynchronization therapy; Digital health; Literature search
    DOI:  https://doi.org/10.1093/ehjdh/ztag125
  4. Cureus. 2026 Jul;18(7): e113428
      Artificial intelligence (AI) holds great promise for improving systematic reviews, particularly as the number of reviews keeps growing. A solution to keep systematic reviews up to date are the so-called living systematic reviews. Although they could be a solution, they currently face challenges regarding everything having to be done manually. A potential solution could be AI-driven reviews that automate each step, from literature searching and translation to data extraction and analysis. Potentially, AI could search through large amounts of studies, sort them by relevance, and extract relevant data. This significantly reduces the workload for researchers, and thereby gives researchers time to focus on oversight rather than manual tasks. However, these methods raise concerns about algorithmic bias and transparency. Proper training, clear ethical guidelines, and interdisciplinary collaboration are keys to ensuring quality and integrity. AI-driven reviews may become essential for efficiently handling the expanding scientific literature.
    Keywords:  artificial intelligence in medicine; living systematic reviews; medical education; research methodology and ethics; systematic reviews and meta-analyses
    DOI:  https://doi.org/10.7759/cureus.113428
  5. Crit Care Explor. 2026 Sep 01. 8(9): e1474
       IMPORTANCE: Large language models (LLMs) are increasingly used for scientific literature retrieval, yet their citation accuracy in specialized clinical domains remains poorly characterized. In neurocritical care (NCC), fabricated or inaccurate citations may be difficult to detect without deliberate verification.
    OBJECTIVES: To evaluate hallucination and fabrication rates of peer-reviewed citations generated by three LLMs across core NCC topics, under constrained zero-shot, memory-only conditions.
    DESIGN, SETTING, AND PARTICIPANTS: In this cross-sectional, blinded technology performance evaluation, Generative Pretrained Transformer (GPT)-5.3, DeepSeek-V3, and Grok-4 were queried on March 10, 2026, under identical zero-shot, retrieval-disabled web-interface conditions. Ten NCC topics were submitted to each model, and each model generated 10 references per topic, yielding 300 references.
    MAIN OUTCOMES AND MEASURES: Two NCC experts, blinded to model identity, independently verified each reference against PubMed, DOI, Google Scholar, and CrossRef and scored accuracy using a Hallucination Scale (0-3). The primary outcome was any hallucination, defined as any citation inaccuracy. The secondary outcome was fabrication, defined as a nonexisting complete bibliographic entity.
    RESULTS: Inter-rater agreement was excellent (κ = 0.91; 95% CI, 0.86-0.96). Overall, 165 of 300 references (55.0%) contained a citation inaccuracy, and 85 of 300 (28.3%) were completely fabricated. DeepSeek-V3 had the lowest hallucination rate (23%; fabrication 8%), followed by GPT-5.3 (69%; fabrication 27%) and Grok-4 (73%; fabrication 50%). Compared with DeepSeek-V3, Grok-4 was 3.17 times more likely to hallucinate (95% CI, 2.03-4.96; p < 0.001), and GPT-5.3 was 3.00 times more likely to hallucinate (95% CI, 1.94-4.63; p < 0.001). Topic-level findings were exploratory and should be interpreted cautiously.
    CONCLUSIONS AND RELEVANCE: Under standardized zero-shot, retrieval-disabled web-interface conditions, LLMs generated substantial numbers of inaccurate and fabricated NCC citations. Because fabricated references can appear complete and credible, artificial intelligence-generated citations should be verified across reliable databases before use in clinical, educational, or scholarly work.
    Keywords:  artificial intelligence; citation hallucination; fabricated references; large language models; neurocritical care
    DOI:  https://doi.org/10.1097/CCE.0000000000001474
  6. Diagnostics (Basel). 2026 Aug 18. pii: 2620. [Epub ahead of print]16(16):
      Background/Objectives: Rosacea is a chronic inflammatory skin disease that requires long-term management and continuous patient education regarding triggers, skincare practices, and treatment adherence. In recent years, patients have increasingly turned to online platforms and artificial intelligence (AI)-based chatbots for health-related information. Although ChatGPT has been evaluated in the context of rosacea, evidence regarding the performance of other AI chatbots remains limited. This study aimed to evaluate the accuracy, reliability, quality, readability, and diagnostic relevance of AI-generated responses to common rosacea-related patient questions and to assess their potential role as sources of health-related information. Methods: Between 21 December 2025 and 22 February 2026, rosacea-related questions were collected from the publicly accessible Quora platform using a systematic screening process. Twenty clinically relevant and representative questions covering diagnosis, triggers, treatment options, skincare practices, and disease manifestations were selected. Each question was independently submitted to three AI chatbots (Claude Sonnet 4.5, ChatGPT-5.2, and DeepSeek-V3.2). Responses were evaluated by domain experts using the modified DISCERN (mDISCERN) for reliability, the Global Quality Scale (GQS) for overall quality, the Flesch Reading Ease Score (FRES) for readability, and a 5-point Likert scale for accuracy. Reference hallucinations were assessed through manual verification of cited sources. Statistical comparisons were performed using the Friedman test with Bonferroni-adjusted post hoc analyses, and effect sizes were calculated using Kendall's coefficient of concordance (Kendall's W). Results: Significant differences were observed among the AI chatbots across all evaluation domains (p < 0.05), with moderate to large effect sizes (Kendall's W = 0.272-0.683). ChatGPT-5.2 and DeepSeek-V3.2 demonstrated significantly higher reliability and accuracy scores than Claude Sonnet 4.5. DeepSeek-V3.2 achieved the highest overall quality scores, whereas ChatGPT-5.2 produced the most readable responses. Reference analysis revealed variable hallucination rates among the evaluated models. Conclusions: Generative AI chatbots demonstrate considerable potential as sources of health-related information for rosacea-related queries. However, variability in performance and reference hallucination rates highlights the need for careful validation before their widespread use as complementary sources of patient health information.
    Keywords:  clinical decision support; diagnostic information; generative artificial intelligence; large language models; rosacea
    DOI:  https://doi.org/10.3390/diagnostics16162620
  7. Front Psychiatry. 2026 ;17 1881536
       Background: Postpartum depression (PPD) is a common perinatal psychiatric disorder with significant implications for maternal and infant health. Artificial intelligence (AI) chatbots have emerged as widely accessible tools for health information, but their performance in providing accurate, reliable, and readable PPD related information remains underexplored.
    Methods: We evaluated six AI chatbots, ChatGPT-5, ChatGPT-4o, Claude Sonnet 4.5, DeepSeek-V3.2, DeepSeek-R1, and Gemini 2.5 Pro, using 200 standardized multiple choice questions (MCQs) on PPD to assess validity. ChatGPT-4o and DeepSeek-R1 were included only in the MCQ based validity analysis as earlier version comparators. Reliability and readability were further assessed using 20 core public education questions in the four latest models: ChatGPT-5, Claude Sonnet 4.5, DeepSeek-V3.2, and Gemini 2.5 Pro. Chatbot performance was assessed across three dimensions: validity, reliability, and readability. Each MCQ was presented three times independently, and each core public education question was assessed once per model.
    Results: In the six model MCQ based validity analysis, ChatGPT-5 achieved the highest overall accuracy on MCQs (97.50% ± 0.50%). In the four models reliability and readability analyses, ChatGPT-5 obtained the highest DISCERN, EQIP, and GQS scores, suggesting relatively better content quality and user oriented usefulness. However, JAMA benchmark scores were low across all models, including ChatGPT-5, indicating limited transparency, source attribution, currency, and disclosure. All models produced outputs exceeding the recommended sixth grade reading level, although ChatGPT-5 and Gemini 2.5 Pro were relatively more accessible.
    Conclusion: AI chatbots, particularly ChatGPT-5, showed potential as supplementary tools for providing postpartum depression related information, especially in standardized MCQ based assessment. However, this study did not evaluate clinical safety, patient comprehension, user behavior, or real world effectiveness, and suboptimal readability may limit accessibility for users with lower health or digital literacy. Inadequate transparency, limited source attribution, and suboptimal readability indicate that AI chatbots should not be used as autonomous sources of postpartum mental health guidance and should not replace professional assessment or care.
    Keywords:  artificial intelligence; health information; large language models; patient education; postpartum depression
    DOI:  https://doi.org/10.3389/fpsyt.2026.1881536
  8. Work. 2026 Aug 26. 10519815261477742
      BackgroundToilet training is a critical developmental milestone that may have long-term implications for children's psychosocial development if improperly guided. With the increasing use of artificial intelligence (AI) chatbot applications for health-related information, evaluating the reliability and readability of AI-generated guidance on sensitive developmental topics has become essential.ObjectiveThis study aimed to evaluate the reliability and readability of responses provided by AI chatbot applications regarding commonly asked questions about toilet training.MethodsTwo widely used AI platforms-ChatGPT-4 Turbo (OpenAI) and Gemini 2.0 Flash (Google)-were included in the study. The study was initiated on April 29, 2025, by submitting a standardized prompt to the chatbots. Ten frequently asked questions about toilet training were selected based on AI-generated query lists and Google Trends data. Responses were obtained in independent sessions and evaluated by a panel of five child development experts using a four-point Likert-type scale developed by Mika et al. Readability levels were assessed using the Flesch-Kincaid Grade Level through WordCalc software. Statistical analyses were conducted to compare quality and readability across platforms.ResultsStatistically significant differences in response quality were identified for Questions 2, 3, and 7 (p < 0.05). In the quality rating system, lower scores indicate higher response quality. Gemini demonstrated lower (better) median quality scores for these items. Regarding readability, Gemini produced responses with a higher Flesch-Kincaid Grade Level (i.e., more complex reading level), particularly for Questions 3 and 7. No statistically significant differences in response quality were found for the remaining seven questions.ConclusionsBoth AI applications provided generally acceptable expert-rated responses to common toilet-training questions; however, differences in response quality and readability were observed across specific items. These findings suggest that AI tools may serve as accessible supplementary informational resources for families, but they should not be interpreted as substitutes for professional guidance or as evidence of clinical validity.
    Keywords:  artificial intelligence; child rearing; comprehension; consumer health information; data accuracy; generative artificial intelligence; health communication; toilet training
    DOI:  https://doi.org/10.1177/10519815261477742
  9. Front Public Health. 2026 ;14 1886493
       Background: Artificial intelligence (AI) has garnered significant attention and, to some extent, has even replaced certain search engines as a popular channel for people across regions to access information. HIV/AIDS is a common topic among infectious diseases, and patients or high-risk groups frequently search for information online. Our study evaluated the accuracy of two AI models (ChatGPT-3.5 and DeepSeek-R1) in answering HIV/AIDS-related questions.
    Methods: This was a cross-sectional analytical study comparing two advanced large language models representing Eastern and Western cultural backgrounds, respectively, with a gold standard (an infectious disease specialist). We compiled 26 common questions related to HIV/AIDS and categorized them into four themes (basic knowledge, diagnosis, treatment, and prevention). Three infectious disease experts independently scored the AI models' responses on a 4-point scale for each answer. There was no statistically significant difference in the scores for HIV/AIDS-related questions between the two AI models: Z = -2.135, two-sided p-value = 0.451 (p > 0.05). DeepSeek-R1 demonstrated higher accuracy, with 53.85% of its responses rated as "Excellent," compared to 48.72% for ChatGPT-3.5; there was no significant difference in accuracy between the two AI models when answering HIV/AIDS-related questions (χ 2 = 0.429, df = 1, p = 0.513). Both AI models achieved high average composite scores (DeepSeek-R1 = 3.50, ChatGPT-3.5 = 3.44, out of a maximum of 4 points). The AI models performed exceptionally well across all domains, though they were relatively weaker in the "Treatment" domain.
    Conclusion: Our study has identified the potential of AI models, particularly ChatGPT-3.5 and DeepSeek-R1, to provide accurate and comprehensive responses to HIV/AIDS-related questions. As free online tools, they can provide personalized, useful medical information to diverse populations across regions. However, medical information generated by AI models should still be used under the supervision and review of healthcare professionals.
    Keywords:  AIDS; ChatGPT-3.5; Deepseek-R1; HIV; health knowledge
    DOI:  https://doi.org/10.3389/fpubh.2026.1886493
  10. J Robot Surg. 2026 Aug 28. pii: 896. [Epub ahead of print]20(1):
      Robot-assisted colorectal cancer resection (RACR) has emerged as a preferred surgical approach for mid-to-low rectal tumors, yet patients often struggle to find accessible, accurate information about this complex procedure. Artificial intelligence (AI) chatbots increasingly serve as lay-friendly health information sources, but whether they deliver counseling of sufficient quality, transparency, and readability for RACR-specific queries remains unknown. To evaluate the quality, transparency, and readability of patient counseling information generated by four leading AI chatbots-ChatGPT (GPT-5.5), Google Gemini (Gemini 3.1 Pro), Anthropic Claude (Claude Sonnet 5), and DeepSeek (DeepSeek-V4)-in response to common questions about robot-assisted colorectal cancer resection. Thirty frequently asked patient questions spanning five clinical domains were posed to each chatbot on a single day (June 20, 2026), yielding 120 question-paired responses. Two independent raters assessed quality using the DISCERN instrument and the Patient Education Materials Assessment Tool (PEMAT). Transparency was measured with a 5-item checklist. Readability was computed via four validated formulas (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index, SMOG), with three syllable-free metrics (Coleman-Liau, Automated Readability Index, Dale-Chall) as a sensitivity analysis. Because the same 30 questions were posed to every chatbot, between-platform comparisons used the Friedman test with question as the blocking factor, with Bonferroni-corrected pairwise Wilcoxon signed-rank tests and Kendall's W effect sizes. Inter-rater reliability was estimated using intraclass correlation coefficients (ICC). ChatGPT achieved the highest mean DISCERN score (64.7 ± 7.0), the only platform whose mean reached the high-quality threshold (≥ 63), followed by Claude (61.2 ± 7.5), Gemini (59.4 ± 7.7), and DeepSeek (56.8 ± 8.2) (Friedman χ²(3) = 13.62, p = 0.003, Kendall's W = 0.15) (Fig. 1); after Bonferroni-corrected paired comparisons, only the ChatGPT-DeepSeek difference remained significant (p = 0.004). Transparency was uniformly deficient: only 27.5% of responses cited sources, and a mere 5.0% indicated information currency (Fig. 3). All chatbots produced text far exceeding the NIH-recommended sixth-grade level (mean Flesch-Kincaid Grade Level 12.3; platform means 11.4-13.7) (Fig. 4). DeepSeek yielded the most readable output (Flesch Reading Ease: 42.1), while ChatGPT generated the most complex text (Flesch Reading Ease: 26.4). A significant negative correlation between quality and readability was observed (ρ = -0.29, p = 0.001) (Fig. 5). AI chatbots furnish moderately good-quality counseling for RACR, yet persistent gaps in transparency and readability render them inadequate as standalone patient education tools. Between-platform quality differences were modest and, once the paired design and multiple comparisons were accounted for, remained significant only between ChatGPT and DeepSeek. No platform met recommended readability benchmarks. Because factual accuracy was not evaluated, no conclusion on clinical safety can be drawn. Clinicians should guide patients toward verified resources and consider supplying optimized prompt templates to enhance accessibility of AI-generated content.
    Keywords:  Artificial intelligence; Chatbot; Colorectal cancer; DISCERN; Large language model; Patient counseling; Readability; Robotic surgery
    DOI:  https://doi.org/10.1007/s11701-026-03874-9
  11. Oral Health Prev Dent. 2026 Aug 26. 24 659-667
       BACKGROUND: As society is increasingly depending on large language models (LLMs) for health-related questions, it is essential to objectively evaluate the quality and accessibility of the oral cancer information they provide. Although LLMs occupy a growing space in digital health communication, it remains unknown whether the information they generate is both reliable and easy to read.
    OBJECTIVE: The purpose of this study was to evaluate the reliability and readability, respectively, of the responses generated by four mainstream LLMs (ChatGPT, Gemini, Perplexity, and DeepSeek) to questions related to oral cancer. Specifically, the present authors aimed to evaluate the reliability of responses to common oral cancer questions and assess whether the readability of responses meets established expectations.
    METHODS: Twenty-two commonly asked, patient-orientated oral cancer-related questions were developed through two predefined phases: Google Trends analysis and expert consultation with specialists in oral oncology. Each question was entered as an independent single-turn prompt into four LLMs: ChatGPT-5, Gemini 2.5, Perplexity Pro, and DeepSeek v3.2. The primary outcome was information reliability and quality, assessed using four standardized instruments: the DISCERN questionnaire, the Ensuring Quality Information for Patients (EQIP) tool, the Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Scale (GQS). The secondary outcome was readability, assessed using six established indices: the Automated Readability Index, Flesch Reading Ease Score, Gunning Fog Index, Flesch-Kincaid Grade Level, Coleman-Liau Index, and Simple Measure of Gobbledygook.
    RESULTS: Significant differences were observed among the four LLMs in DISCERN, EQIP, and JAMA scores (all P 0.001), whereas no significant difference was found in GQS scores (P = 0.440). Perplexity Pro achieved the highest mean DISCERN score (46.36 ± 4.70), EQIP score (85.00 ± 0.00), GQS score (4.05 ± 0.58), and JAMA score (1.00 ± 0.00). However, all models produced responses above the recommended sixth-grade readability level. The mean FKGL scores ranged from 12.65 ± 3.07 for ChatGPT-5 to 15.65 ± 3.36 for Perplexity Pro, and the mean FRES scores ranged from 37.50 ± 14.27 for Perplexity Pro to 52.59 ± 12.47 for Gemini 2.5.
    CONCLUSION: Current LLMs may support oral cancer patient education, but their use remains limited by variable information quality, insufficient transparency, and poor readability. Although Perplexity Pro performed better on several reliability-related metrics, no model showed consistently high performance across all dimensions or met recommended readability standards. Future LLM-based patient education tools should prioritise verifiable sourcing, guideline-based accuracy, risk communication, and plain-language adaptation.
    Keywords:  AI; chatbot; health information; oral cancer; readability; reliability
    DOI:  https://doi.org/10.3290/j.ohpd.c_2786
  12. Digit Health. 2026 Jan-Dec;12:12 20552076261475785
       Background: Artificial intelligence (AI) chatbots are increasingly used for online health information seeking and have become a common source of information on postherpetic neuralgia (PHN). The quality and readability of the information they provide have not been systematically evaluated.
    Objective: To assess the quality, transparency, and readability of postherpetic neuralgia information generated by mainstream artificial intelligence chatbots and to examine whether this content meets the recommended sixth-grade reading level for patient education materials.
    Methods: Using the MeSH term Neuralgia, Postherpetic, we identified 11 highly relevant postherpetic neuralgia-related search queries worldwide from 2020 to 2025 through Google Trends. These queries were entered verbatim into ChatGPT-4o, Gemini-1.5, Perplexity Pro, and Copilot using a standardised input procedure. Responses were assessed with DISCERN, EQIP, JAMA, and GQS for information quality, completeness, transparency, and overall educational value. Readability was evaluated with ARI, FRES, GFI, FKGL, CL, and SMOG. The Kruskal-Wallis test was used for between-model comparisons, and the Wilcoxon signed-rank test was used to compare readability indices against sixth-grade reading benchmarks.
    Results: DISCERN, EQIP, and GQS scores differed significantly across chatbots in DISCERN (P < 0.001), EQIP (P < 0.001), and GQS (P < 0.001) scores among the four chatbots; JAMA scores showed no significant difference (P = 0.668). No chatbot reached an excellent overall rating, although Perplexity Pro and Copilot performed better than ChatGPT-4o and Gemini-1.5 on most quality-related measures. For readability, none of the chatbots met the sixth-grade reading standard on any metric: all FRES scores were <80, and all other indices (ARI, GFI, CL, FKGL, SMOG) exceeded the grade-6 thresholds. Readability did not differ significantly among chatbots (all P > 0.2).
    Conclusion: Mainstream artificial intelligence chatbots can provide postherpetic neuralgia information with moderate quality, but the readability of this content remains consistently above recommended patient education levels. The findings should be interpreted within the scope of the selected chatbots, default operational settings, search queries, and the time-sensitive nature of chatbot outputs. Further improvement in source reporting and plain-language presentation remains warranted.
    Keywords:  AI chatbot; postherpetic neuralgia; quality; readability
    DOI:  https://doi.org/10.1177/20552076261475785
  13. Digit Health. 2026 Jan-Dec;12:12 20552076261481386
       Background: Patients are increasingly turning to online resources and artificial intelligence (AI)-based tools to obtain information about orthopedic conditions and surgical options. Large language models, such as ChatGPT, are becoming prominent in patient education; however, their reliability and readability remain uncertain. This study evaluated the quality and readability of responses generated by ChatGPT-4o and ChatGPT-5 to frequently asked patient questions regarding hallux rigidus fusion surgery.
    Methods: Twenty commonly asked patient questions were compiled and presented to ChatGPT-4o and ChatGPT-5. Readability was assessed using the Flesch-Kincaid Grade Level, Gunning Fog, Coleman-Liau, and Simple Measure of Gobbledygook indices. Quality was evaluated with the DISCERN tool, response accuracy scores, and Journal of the American Medical Association (JAMA) criteria. Interrater agreement was measured using the Intraclass Correlation Coefficient (ICC).
    Results: ChatGPT-4o generated longer responses (802 vs. 242 words; p<0.001) with slightly higher readability grade levels (10.81 vs. 10.37; p=0.031). Accuracy (2.00 vs. 1.85; p=0.323) and DISCERN scores (49.35 vs. 48.93; p=0.747) showed no significant differences. All responses received a JAMA score of 0 due to the absence of citations, authorship, or transparency indicators. Interrater reliability indicated moderate to good agreement (ICC: 0.68-0.80).
    Conclusion: ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery. Although linguistically coherent, responses lack evidence-based detail and individualized guidance. These models may supplement, but cannot replace, expert orthopedic counseling. Ensuring physician oversight and integrating validated, updated clinical content remain essential for safe implementation of AI-generated patient information.
    Keywords:  ChatGPT; arthrodesis; artificial intelligence; frequently asked questions; hallux rigidus; readability
    DOI:  https://doi.org/10.1177/20552076261481386
  14. Laryngoscope Investig Otolaryngol. 2026 Oct;11(5): e70526
       Objective: Evaluate the accuracy, comprehensiveness, and similarity to provider response of ChatGPT-4o in providing patient education regarding microtia and aural atresia management.
    Methods: Ten standardized inquiries were created across three domains (General Information, Hearing/Anatomy, Treatment Decision-Making) and entered into ChatGPT-4o. Responses were evaluated by an international multidisciplinary cohort of 23 experts, including pediatric otolaryngologists and audiologists. Using 5-point Likert scales, graders assessed factual accuracy, comprehensiveness, similarity to physician responses, and overall quality. Readability was analyzed using Flesch-Kincaid metrics.
    Results: ChatGPT-generated responses demonstrated high accuracy (mean = 4.03 [SD = 0.66]), comprehensiveness (mean = 4.18 [SD = 0.54]), and similarity to the provider's response (4.02 [0.70]). Audiologists assigned significantly lower ratings than pediatric otolaryngologists regarding accuracy (mean = 3.76 vs. 4.13; p < 0.001) and comprehensiveness (3.94 vs. 4.30; p = 0.046). Qualitative analysis included concerns regarding advanced terminology, omission of specific risks, and a lack of empathetic framing regarding parental guilt. Flesch-Kincaid Reading Grade Level often exceeded the 5th-grade average level for patient education materials (mean = 9.64 [SD = 1.21]).
    Conclusion: ChatGPT-4o provides reliable introductory information but lacks the technical nuance and multidisciplinary coordination essential for complex microtia care. While a useful adjunct for early information-seeking, ChatGPT-4o lacks the multidisciplinary framing and depth required for ideal counseling.
    Keywords:  ChatGPT; microtia; microtia reconstruction; patient information
    DOI:  https://doi.org/10.1002/lio2.70526
  15. Front Public Health. 2026 ;14 1924162
       Background: Generative artificial intelligence chatbots are increasingly used as sources of consumer health information. Although their performance has been examined in several medical conditions, evidence specific to hand, foot, and mouth disease (HFMD) remains limited.
    Objective: To compare the safety, accuracy, empathy, information quality, and readability of HFMD-related responses generated by five publicly accessible chatbots.
    Methods: In this exploratory cross-sectional study, 20 researcher-developed English-language prompts were constructed from authoritative public-health sources, Google Trends topic mapping, and caregiver-informed wording refinement. ChatGPT-4o, Gemini 2.5 Pro, Copilot, Doubao, and DeepSeek-V3.2-Exp were evaluated between April 2 and April 5, 2026. Five trained reviewers independently assessed responses against predefined reference standards using safety and accuracy criteria, an empathy scale, DISCERN, EQIP, the Global Quality Score, and JAMA benchmarks. Six formula-based readability indices were calculated. Paired comparisons used Friedman and Cochran Q tests, with prespecified post-hoc procedures and Benjamini-Hochberg correction.
    Results: Unsafe-response rates ranged from 5.0 to 15.0%, with no statistically significant difference detected among chatbots (p = 0.797). No statistically significant inter-model differences were detected for accuracy, empathy, DISCERN, EQIP, JAMA, or Global Quality Score in the 20-prompt set; because the study was exploratory and was not powered to establish equivalence, these findings do not demonstrate comparable or interchangeable performance. All six readability indices differed significantly among chatbots (p = 0.030 to <0.001). ChatGPT and Doubao generally produced lower estimated grade-level complexity than Gemini and DeepSeek. Ten potentially unsafe or misleading responses were identified, mainly involving overgeneralization of EV71 vaccine protection, hand-hygiene qualification, and disinfection advice.
    Conclusion: Across a limited set of standardized English-language prompts, the five chatbots often generated coherent HFMD information, but occasional safety-relevant inaccuracies and readability barriers remained. These systems may support general information seeking, but their responses require cautious interpretation and should not replace individualized professional advice.
    Keywords:  HFMD; and mouth disease; artificial intelligence; chatbot; foot; hand; health communication; large language model; pediatric infectious disease; readability
    DOI:  https://doi.org/10.3389/fpubh.2026.1924162
  16. J Vis Exp. 2026 Aug 21.
      Large language models (LLMs) are increasingly used for patient health education, yet the readability and educational quality of LLM-generated information on trigeminal neuralgia (TN) have been insufficiently evaluated. This cross-sectional benchmarking study compared TN educational content generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Twenty frequently asked TN questions covering basic disease knowledge, etiology/risk factors, diagnosis, treatment, and prevention/rehabilitation were presented to each model using standardized prompts. Readability was assessed using seven established indices, educational suitability using the Patient Education Materials Assessment Tool for Understandability and Actionability (PEMAT), and overall information quality using the Global Quality Score (GQS). Two clinical experts independently evaluated all responses, with disagreements resolved by a senior adjudicator. Statistical analyses compared model performance, thematic differences, and correlations among the evaluation metrics. Significant differences were observed among the models for readability, PEMAT, and GQS scores. GPT-5 generated the most linguistically complex responses but achieved the highest ratings for educational suitability and information quality. In contrast, Wenxin Yiyan produced the most readable text but generally scored lower on PEMAT and GQS. Content category influenced readability, with prevention/rehabilitation and etiology/risk-factor topics being more difficult to read, whereas PEMAT and GQS remained relatively consistent across themes. Readability indices showed strong internal consistency and weak-to-moderate positive correlations with PEMAT and GQS, while PEMAT and GQS demonstrated a moderate positive correlation. These findings suggest that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability. Because patient comprehension, satisfaction, trust, health outcomes, factual accuracy, and clinical safety were not evaluated, these results should be interpreted as an expert-rated benchmarking analysis rather than evidence of clinical readiness.
    DOI:  https://doi.org/10.3791/70833
  17. Clin Ter. 2026 Sep-Oct;177(5):177(5): 1216-1233
       Background: Artificial intelligence chatbots, particularly ChatGPT, have emerged as increasingly popular sources of health information for the general public. However, concerns persist regarding the accuracy, safety, and appropriateness of AI-generated medical advice, especially for life-threatening conditions such as myocardial infarction.
    Objectives: This study aimed to evaluate the accuracy, completeness, and safety of ChatGPT-generated responses and information to public questions about heart attacks, assess user perceptions and trust, and compare AI-generated content with expert-validated sources.
    Methods: A cross-sectional study was conducted between January and March 2026. A total of 73 commonly asked heart attack-related questions were submitted to ChatGPT (GPT-4), and responses were independently evaluated by two expert reviewers and one external reviewer using standardized rubric assessing accuracy, completeness, clarity, safety, and tone. Inter-rater reliability was assessed using intraclass correlation coefficients. In parallel, a survey of 352 participants evaluated user perceptions, trust, and behavioral use of ChatGPT for medical information. Statistical analyses included descriptive statistics, group comparisons, correlation analyses, and multivariable regression models to identify predictors of trust and satisfaction.
    Results: Expert reviewers assigned high ratings for accuracy (mean: 4.62 ± 0.62 and 4.48 ± 0.63) and safety (mean: 4.78 ± 0.51 and 4.05 ± 0.50), with substantial inter-rater agreement (ICC = 0.788). In contrast, the external reviewer documented considerably lower completeness ratings (2.59 ± 0.57), revealing deficiencies in information coverage. Remarkably, 40.6% of respondents indicated they had followed ChatGPT's medical recommendations without seeking professional consultation, while 17.6% reported depending on the chatbot during urgent medical situations. Levels of user trust and satisfaction were moderate, showing significant correlations with perceived accuracy, usage frequency, and healthcare-related educational background.
    Conclusions: Expert assessments confirmed that ChatGPT generally provides accurate and safe cardiovascular health information; nevertheless, considerable inter-reviewer variability-especially the markedly lower completeness scores from the external evaluator-reveals significant gaps in content depth. These results indicate that ChatGPT functions more appropriately as an adjunctive educational resource rather than a replacement for professional medical consultatio n, emphasizing the necessity for enhanced safety mechanisms, clearer user instructions, and standardized content protocols when addressing high-stakes health topics. ChatGPT should be considered an adjunctive resource rather than a substitute for urgent professional medical assessment in suspected heart attack cases.
    Keywords:  ChatGPT; artificial intelligence; digital health; health literacy; heart attack; myocardial infarction; patient education
    DOI:  https://doi.org/10.7417/CT.2026.2124
  18. Front Public Health. 2026 ;14 1905883
       Background: Patients increasingly use public large language model chatbot interfaces to seek health information. In vascular disease and perioperative management, default first responses may influence how patients interpret urgent symptoms, antithrombotic medications, procedural choices, and anesthesia-related safety issues.
    Methods: This cross-sectional benchmark study evaluated 110 default first responses returned by five publicly accessible LLM chatbot interfaces during a defined access window on May 28-29, 2026, Beijing time (UTC + 8). Interface names were recorded solely as the public-interface display labels visible at the time of access and should not be interpreted as independently verified API-level model identifiers. Each model was queried with 22 guideline-derived English patient-facing questions, yielding 110 responses. Responses were assessed using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), Journal of the American Medical Association (JAMA) benchmark criteria, used only as a visible metadata/transparency proxy, six readability formulas, an investigator-developed Guideline Concordance Score, and an investigator-developed Potential Clinical-Risk Severity Flag. Differences across the evaluated public-interface response sets were tested using Friedman tests with Holm-adjusted post hoc comparisons.
    Results: Interrater agreement was high for established instruments: DISCERN ICC(A,1) = 0.940, EQIP ICC(A,1) = 0.830, GQS weighted κ = 0.829, and JAMA weighted κ = 0.898. Observed public-interface response performance differed significantly for DISCERN, EQIP, and JAMA criteria (all p < 0.001), and for GQS (p = 0.011). Within this defined late-May 2026 response set, responses returned by the interface displaying the label 'Grok-4.3' had the highest observed sample means for DISCERN (58.68), EQIP (81.14), GQS (4.27), and the JAMA visible metadata/transparency proxy score (1.00). No evaluated public-interface response set achieved recommended sixth-grade readability, and no individual response met all six readability thresholds. Responses returned by the interface displaying the label 'DeepSeek-v4' had the most favorable observed readability profile within this sampled response set, although the mean FKGL remained 11.53. These observations should not be interpreted as durable model rankings.
    Conclusion: In this English-language benchmark of default first responses from five public LLM interfaces accessed through specific logged-in accounts from a Hong Kong, China IP address during a single late-May 2026 window, 107 of 110 responses received the maximum Guideline Concordance Score, and no response met the prespecified criteria for a high Potential Clinical-Risk Severity Flag. Pronounced ceiling and floor effects preclude conclusions about clinical sufficiency or safety. These findings should not be generalized to non-English use, different health-literacy levels, country-specific emergency-care pathways, other regions or account configurations, or later interface states.
    Keywords:  clinical risk; large language models; patient education; perioperative management; readability; vascular disease
    DOI:  https://doi.org/10.3389/fpubh.2026.1905883
  19. Ann R Coll Surg Engl. 2026 Aug 27.
       INTRODUCTION: Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO.
    METHODS: A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores.
    RESULTS: ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores.
    CONCLUSIONS: There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.
    Keywords:  Artificial intelligence; Generative AI; Large language models; Patient education; Peri-acetabular osteotomy
    DOI:  https://doi.org/10.1308/rcsann.2026.0056
  20. Mil Med. 2026 Aug 27. pii: usag394. [Epub ahead of print]
       INTRODUCTION: Artificial intelligence (AI) tools are increasingly used by patients and caregivers seeking medical information. Developmental dysplasia of the hip (DDH) is a common pediatric condition that frequently prompts parental questions regarding screening, diagnosis, and treatment. The purpose of this study was to evaluate the accuracy, completeness, and clarity of responses generated by a Department of Defense-approved AI chatbot (NIPR-GPT) when answering commonly asked DDH-related questions.
    METHODS: Twelve frequently asked patient questions related to DDH were submitted to the NIPR-GPT chatbot. Each question was entered into a new chatbot session to minimize contextual bias. AI-generated responses were independently evaluated by 5 fellowship-trained pediatric orthopedic surgeons. Responses were graded using a 4-point scale assessing accuracy, completeness, and clarity: (1) unsatisfactory (major inaccuracies requiring substantial correction), (2) satisfactory with moderate clarification required, (3) satisfactory with minimal clarification required, and (4) excellent with no clarification required. Inter-rater reliability was assessed using intraclass correlation coefficients (ICC) and weighted kappa statistics, while overall internal consistency among raters was assessed using Krippendorff's alpha.
    RESULTS: All AI-generated responses (100%) were rated either satisfactory or excellent. Six responses (50%) were graded excellent, requiring no clarification, and 6 responses (50%) were graded satisfactory with minimal clarification required. No responses were graded as unsatisfactory or requiring major correction. Single-rater reliability among individual reviewers was poor (ICC [2, 1] = 0.12), reflecting variability in reviewer thresholds. However, reliability improved when scores were averaged across raters, demonstrating moderate agreement (ICC [2, k] = 0.41). Internal consistency among reviewers was Krippendorff's α = 0.163, indicating heterogeneity in reviewer grading but not systematic deficiencies in AI responses. Five fellowship-trained pediatric orthopedic surgeons independently graded all 12 responses. No reviewer graded any response as unsatisfactory or requiring substantial correction. Mean question scores ranged from 2.8 to 4.0 across the 12 questions, with an overall mean score of approximately 3.6.
    CONCLUSION: NIPR-GPT generated generally accurate, clear, and clinically appropriate responses to common DDH-related questions. Half of responses required no clarification, and none required major correction. These findings suggest that AI chatbots may serve as useful adjunct educational tools for patients and caregivers within military healthcare systems. However, variability in expert interpretation and the potential for patient misinterpretation underscore the continued importance of clinician oversight when integrating AI-generated information into patient education.
    DOI:  https://doi.org/10.1093/milmed/usag394
  21. Ear Nose Throat J. 2026 Aug 22. 1455613261481553
      BackgroundLaryngomalacia is the most common cause of stridor in infants. Guidelines recommend parent education materials be written at a sixth to eighth-grade reading level, yet compliance of online resources remains unclear.AimTo evaluate the readability of online parental education materials about laryngomalacia across multiple sources and geographic regions.MethodsA systematic search identified the top 10 Google results, hospital resources from Australia (n=2), the UK (n=2), and the US (n=5), and outputs from five AI chatbots. Readability was assessed using Flesch Reading Ease Score (FRES), Flesch-Kincaid Grade Level (FKGL), and Fry Grade Level. Kruskal-Wallis tests compared readability across source types and regions.ResultsOf 19 unique resources, only one (5.3%) met recommended reading levels. Mean FKGL was 10.61 (SD=2.24) across all resources, significantly exceeding guidelines. Australian hospitals were most difficult (mean FKGL=13.80); UK hospitals were most accessible (mean FKGL=9.45), though still exceeding recommendations. Significant regional differences existed for FRES (H=6.43, p=0.038) and Fry Grade Level (H=6.19, p=0.045), but not FKGL (H=4.13, p=0.127).Conclusions94.7% of laryngomalacia parental education materials exceeded recommended readability levels. AI chatbots and Australian hospital websites were particularly inaccessible. Healthcare providers should prioritise simplifying laryngomalacia education materials to support health literacy and parental comprehension.
    Keywords:  health literacy; laryngomalacia; otolaryngology; paediatric otolaryngology; parent education; patient education; readability
    DOI:  https://doi.org/10.1177/01455613261481553
  22. Toxins (Basel). 2026 Aug 14. pii: 347. [Epub ahead of print]18(8):
      Mycotoxins contaminate a wide range of foods and may contribute to both acute toxicity and chronic dietary exposure. Because online video increasingly shapes public interpretation of food-safety hazards, we evaluated the quality of English-language YouTube™ content on foodborne mycotoxin risks. Of 166 records screened for eligibility, 53 were excluded and 113 videos were analyzed. Two food-safety experts independently evaluated each video with an investigator-developed Video Content Quality (VCQ) checklist, the Global Quality Scale (GQS), and a five-item modified DISCERN instrument. The final sample comprised 77 company/commercial and 36 academic/noncommercial videos. After Holm adjustment across five source-group outcomes, company/commercial videos had higher VCQ and modified DISCERN scores (both adjusted p < 0.001), whereas academic/noncommercial videos had a higher interaction rate (adjusted p = 0.024). The GQS comparison did not meet the adjusted significance threshold (adjusted p = 0.066) and should not be interpreted as proof of equivalence. Video age was associated with cumulative views and likes but not with any quality score after correction. Item-level VCQ results showed that authoritative-source citation and health-effects coverage were among the least frequently awarded full-credit domains. Although the three quality measures were strongly correlated, neither interaction rate nor viewing rate was significantly associated with them. Platform engagement therefore did not serve as a reliable marker of scientific quality in this sample. Source- and region-based comparisons remain exploratory because the single-query design, broad uploader categories, and uneven source composition may have shaped the sampled videos.
    Keywords:  DISCERN; YouTube™; digital health information; food safety education; food toxicology; mycotoxins; risk communication
    DOI:  https://doi.org/10.3390/toxins18080347
  23. Jpn J Radiol. 2026 Aug 26.
       PURPOSE: To evaluate the quality and accuracy of YouTube videos regarding PET/CT radiation safety and to assess the feasibility of using a Large Language Model (LLM) as an automated tool for content moderation compared to a human expert.
    METHODS: A systematic search was conducted on YouTube using keywords related to "PET/CT radiation safety." A total of 42 videos were included and categorized by uploader source (Professional vs. Non-Professional). Video quality was assessed using the modified DISCERN (mDISCERN) tool and the Global Quality Scale (GQS). Usefulness and engagement metrics (views, likes) were analyzed. Additionally, inter-rater reliability between a human physician with clinical work experience in a nuclear medicine department and an AI model (Gemini-3) was evaluated using Cohen's Kappa (κ).
    RESULTS: The majority of content (90.5%) originated from professional sources. Professional videos demonstrated significantly higher information quality compared to non-professional videos (mean mDISCERN: 3.71 vs. 2.50, P = 0.032; mean GQS: 4.18 vs. 2.50, P = 0.005). However, video popularity was not an indicator of quality, as no significant correlation was found between view counts and mDISCERN scores (ρ = 0.124, P = 0.432). Notably, the agreement between the human expert and the AI model was only slight (κ = 0.191). Discrepancy analysis revealed that the AI model systematically overestimated the quality of content containing commercial bias while penalizing highly technical academic videos for their presentation style.
    CONCLUSION: YouTube contains high-quality information on PET/CT safety, predominantly from professional sources; however, the platform's algorithm does not prioritize clinical accuracy, often favoring lower-quality content. Furthermore, current LLMs lack the contextual judgment required to reliably detect commercial bias and misinformation, underscoring the irreplaceable role of physician oversight in digital patient education.
    Keywords:  Artificial intelligence; PET/CT; Patient education; Radiation safety; Social media; YouTube
    DOI:  https://doi.org/10.1007/s11604-026-02059-6
  24. Asian Pac J Cancer Prev. 2026 Aug 01. pii: 92328. [Epub ahead of print]27(8): 3029-3036
       BACKGROUND: The growing dependence on digital platforms for health-related information has positioned YouTube as a key provider of vaccine-related content in India. Nonetheless, the quality, reliability, and completeness of the information accessible on this platform are still questionable. This study sought to assess the content, quality, and reliability of YouTube videos pertaining to HPV vaccination within the Indian context.
    METHODS: Search terms "Cervical Cancer Vaccination" and "HPV Vaccine" were used to cross-sectionally analyze YouTube videos. Screening the first 200 videos yielded 81 suitable English or Hindi videos. Data on video attributes and engagement was captured. Audio-Visual Quality (AVQ), Global Quality Score (GQS), modified Quality Criteria for Consumer Health Information (mDISCERN), and Video Information and Quality Index assessed quality and dependability. The new Content Score for Cervical Cancer Vaccination (CSCCV) was developed based on national and international guidelines to assess content completeness.
    RESULTS: A total of 81 videos were analyzed and median duration of the videos 169 seconds, Health-related channels accounted for 51(63.0%) of uploads, and doctors or medical experts serving as speakers in 46(56.8%) of cases. Although none of the videos depicted HPV vaccination unfavourably, engagement metrics did not align with the level of information presented. Videos involving medical professionals exhibited markedly higher mDISCERN scores (p = 0.001), signifying enhanced reliability. The overall content completeness was moderate, exhibiting notable deficiencies in guideline-based information.
    CONCLUSION: YouTube videos regarding HPV vaccination in India are predominantly favourable, although they exhibit variability in dependability and comprehensiveness. Enhancing expert-driven, guideline-compliant, and multilingual digital resources is crucial for promoting HPV vaccine adoption and advancing cervical cancer prevention initiatives in India.
    Keywords:  Cervical cancer; HPV vaccination; Health Information Quality; Youtube
    DOI:  https://doi.org/10.31557/APJCP.2026.27.8.3029
  25. Front Digit Health. 2026 ;8 1875817
       Background: Systemic lupus erythematosus (SLE) is a chronic autoimmune disease involving multiple organ systems. Patient education is an essential component of disease management. With the rapid development of digital media, short-video platforms have become increasingly important health information sources for SLE patients. Therefore, the quality of health information disseminated through these platforms warrants attention. However, evidence regarding the quality of SLE-related health information on short-video platforms remains limited.
    Methods: On September 18, 2025, SLE-related videos were retrieved from TikTok, Rednote (Xiaohongshu), and Bilibili. We evaluated the videos' reliability using the DISCERN instrument, their overall quality using the Global Quality Scale (GQS), and their understandability and actionability using the Patient Education Materials Assessment Tool (PEMAT-U/A).
    Results: A total of 276 SLE-related videos from three platforms were analyzed. Significant differences in video quality and reliability were observed among platforms (p < 0.001). Specifically, videos on Rednote exhibited the highest quality, while TikTok videos demonstrated the lowest quality. Videos produced by online health science communicators and certified medical professionals had significantly higher DISCERN and PEMAT-U scores than patient-generated videos (p < 0.05). Correlation analysis indicated that video engagement metrics were positively associated with quality and reliability on Bilibili (p < 0.01), whereas a negative correlation was found on TikTok and Rednote (e.g., DISCERN vs. likes on TikTok: r = -0.327, p = 0.001; GQS vs. collections on Rednote: r = -0.365, p = 0.006). Regarding the relationship between quality and engagement, DISCERN and GQS, and PEMAT-U scores on Bilibili were positively correlated with engagement metrics (p < 0.01). In contrast, several inverse associations were identified on TikTok and Rednote, including negative correlations between DISCERN scores and likes on TikTok (r = -0.327, p = 0.001) and between GQS scores and collections on Rednote (r = -0.365, p = 0.006).
    Conclusion: The quality and reliability of SLE-related videos vary significantly across platforms and are generally suboptimal. The relationship between video quality and user engagement is platform- and measure-specific, suggesting higher-quality health information does not always translate into greater user engagement. Enhancing reliable health information requires coordinated efforts to improve both content quality and audience engagement.
    Keywords:  health information; quality assessment; short video; social media; systemic lupus erythematosus
    DOI:  https://doi.org/10.3389/fdgth.2026.1875817
  26. Front Public Health. 2026 ;14 1820704
       Background: Hypertension is one of the major modifiable cardiovascular risk factors in China. Persistent and poorly controlled hypertension contributes to a wide spectrum of adverse cardiovascular outcomes, yet public awareness and understanding of the condition remain suboptimal. With the rapid expansion of digital media, social platforms have emerged as important channels for disseminating health information. This study evaluates the quality of hypertension-related content presented in short-video formats on widely used social media platforms.
    Methods: Newly registered accounts were used to retrieve the top 100 videos from Douyin, Bilibili, and Kwai based on each platform's default comprehensive ranking algorithm. The collected videos were subsequently screened following predefined criteria, which included removal of duplicates, silent clips, advertisements, irrelevant content, and non-Chinese language videos. In addition, videos produced by the same creator on the same platform were consolidated into a single entry to avoid redundancy. The reliability of the videos was assessed using the DISCERN instrument, and their overall quality was evaluated with the Global Quality Scale (GQS). A comparative analysis was further performed to examine differences in video characteristics and quality across the three short-video platforms.
    Results: A total of 267 videos from various creators were included in the analysis, with Douyin contributing 36.3% (140/267), Bilibili 31.1% (83/267), and Kwai 32.6% (87/267). With respect to video sources, non-specialist physicians produced the largest proportion of content (52.4%, 140/267), followed by specialist physicians (31.5%, 84/267). In the assessment of overall video quality, Bilibili demonstrated the highest completeness scores (p < 0.001). Douyin and Bilibili demonstrated comparable Global Quality Scores (GQS) and both performed significantly better than Kwai (p < 0.001). Nevertheless, all three platforms exhibited uniformly low DISCERN scores. When comparing creators with different professional backgrounds, videos produced by non-profit organizations and specialist physicians showed significantly higher overall quality than those generated by non-specialist physicians and individual users.
    Conclusion: Overall, the quality of the videos was suboptimal. Although hypertension-related content produced by nonprofit organizations and cardiovascular specialists achieved higher scores in overall quality, the presentation of these videos was relatively uniform. This homogeneity limited the breadth of hypertension-related knowledge that viewers could obtain from short-video platforms.
    Keywords:  DISCERN; global quality scale (GQS); health education; hypertension; information quality; short videos
    DOI:  https://doi.org/10.3389/fpubh.2026.1820704
  27. PLOS Digit Health. 2026 Aug;5(8): e0001639
      Monitoring blood glucose is essential for diabetes management, and many patients use short-video platforms to learn capillary blood glucose testing. We evaluated whether previously reported concerns about online health information also apply to this narrower procedural task and compared high-ranking videos on TikTok/Douyin and Bilibili. We included 192 videos (TikTok, n = 100; Bilibili, n = 92) and extracted duration, engagement metrics, uploader type, and video category. Two trained nursing reviewers assessed general educational quality and reliability-related characteristics using the Global Quality Score (GQS) and modified DISCERN (mDISCERN). Between-platform comparisons used Mann-Whitney U tests or Pearson's chi-square tests, with standardized Mann-Whitney effect sizes; Spearman correlations examined associations with quality scores. TikTok videos received more likes, platform-specific saving actions, comments, and shares and were shorter than Bilibili videos (all P < 0.001; r = 0.315-0.403). TikTok also had higher proportions of accounts classified as professional uploaders (63.0% vs 35.9%, P < 0.001) and popular-science videos (53.0% vs 32.6%, P = 0.004). However, GQS and mDISCERN scores did not differ significantly between platforms (GQS, P = 0.246; mDISCERN, P = 0.072). Coverage of four prespecified procedural elements was incomplete on both platforms, with guidance on first-drop discard, alcohol drying, and post-puncture pressing time appearing in only a minority of videos. TikTok achieved substantially greater engagement, but popularity did not indicate significantly better educational quality or reliability-related characteristics. Standardized, professionally developed procedural videos may improve online diabetes self-management guidance.
    DOI:  https://doi.org/10.1371/journal.pdig.0001639
  28. Oxf Open Digit Health. 2026 ;4 oqag019
      Migraine is one of the most common neurological conditions affecting people worldwide. Migraine is the second leading cause of the global burden of neurological diseases. Studies have found the information on TikTok on migraine to be inaccurate, yet it has been a source of health information among the public. The objective of the study was to describe the TikTok videos discussing migraine, to evaluate its content in terms of reliability, quality, and understandability, and investigate the factors of these videos. A search was conducted using the keyword '#migraine' and the top videos were assessed by two independent raters. Video parameters such as type of uploader, views, likes, comments, shares, and duration were also collected. The evaluated videos had low quality and reliability on the Global Quality Scale (GQS) and the modified DISCERN tool, with low understandability and actionability. Higher scores overall in Patient Education Materials Assessment Tool (PEMAT) showed a positive correlation with higher scores on the GQS. Video length (in seconds) had a positive correlation with GQS and PEMAT scores [length with GQS r = 0.393 (P < .001); length with PEMAT Total r = 0.420 (P < .001)]. TikTok videos on migraine were found to be of low quality, reliability, and understandability. These findings highlight the role of social media in contributing to the health illiteracy of its audiences, further lobbying that healthcare institutions should create official social media accounts to spread reliable health information to the public on migraines.
    Keywords:  TikTok; health information; migraine; quality analysis; social media
    DOI:  https://doi.org/10.1093/oodh/oqag019
  29. JMIR Dermatol. 2026 Aug 25. 9 e79250
       Unlabelled: In this cross-sectional analysis of popular melanoma-related TikTok videos, physician-created, educational, and longer videos were associated with higher information quality and audience engagement, although overall DISCERN scores remained low and treatment guidance was frequently absent, highlighting opportunities to improve comprehensive, physician-led melanoma education on social media.
    Keywords:  TikTok; dermatology; melanoma; social media; trends
    DOI:  https://doi.org/10.2196/79250
  30. Health Informatics J. 2026 Jul-Sep;32(3):32(3): 14604582261485143
      ObjectiveThis study aimed to evaluate the quality and reliability of dengue fever-related videos on social media platforms.MethodsA cross-sectional study was conducted on January 10, 2026. Top 120 videos from each platform (YouTube and Douyin) were screened. Video characteristics were extracted. Quality and reliability were assessed using the Global Quality Scale (GQS), modified DISCERN, JAMA Benchmark criteria, and the Content Completeness Score (CCS).Results170 videos (70 from YouTube, 100 from Douyin) were included. YouTube videos were longer (118.50 vs. 78.50, Z = -3.79, p < 0.001), had higher CCS (8.00 vs. 6.00, Z = -3.50, p < 0.001) and mDISCERN scores (3.00 vs. 2.50, Z = -3.74, p < 0.001). YouTube videos were dominated by institutions (91.43%), followed by 8.57% healthcare professionals (HCPs), whereas nearly one third of contributors on Douyin were HCPs (31.00%). Associations between engagement and quality indicators were generally weak and varied across metrics, while engagement indicators on Douyin were strongly intercorrelated.ConclusionsYouTube provided more comprehensive and reliable information, whereas Douyin content was shorter and less comprehensive. Both platforms showed inconsistent alignment between popularity and informational quality, with this disconnect appearing more pronounced on Douyin.
    Keywords:  cross-sectional; dengue fever; douyin (tiktok); social media; youtube
    DOI:  https://doi.org/10.1177/14604582261485143
  31. Geriatr Nurs. 2026 Aug 26. pii: S0197-4572(26)00533-1. [Epub ahead of print]73 104328
      As the amount of health information available on the Internet continues to expand (such as digital health records, online consultation services, and support groups), people have more opportunities to enhance their daily well-being. Meanwhile, the Internet became integral to accessing health information and services during and after COVID-19. The evidence regarding the Internet's impact on improving the well-being of older adults, however, is mixed and inconsistent in academic research. To address these inconsistencies and better understand the Internet's role, this integrative review synthesized literature on the health information behaviour of older adults. Six databases were searched for English-language studies published from the earliest records available in each database through September 2025. After screening, 87 studies were included in the synthesis. This review identifies and explains three major themes: health information needs, health information seeking, and health information use. It demonstrates that although the Internet can be a valuable health information resource for older adults, its usefulness is shaped by age, digital and health literacy, social support, trust, Internet design, and the availability of interpersonal and professional sources. Overall, the Internet is best understood as a supplementary health information source for older adults. These findings can help healthcare providers and web designers develop more accessible, trustworthy, and supportive online tools for older adults in the post-pandemic era.
    Keywords:  Health; Information seeking; Information sources; Internet; Literature review; Older people
    DOI:  https://doi.org/10.1016/j.gerinurse.2026.104328