bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–09–27
thirty papers selected by
Thomas Krichel, Open Library Society



  1. J Can Health Libr Assoc. 2026 Aug;47(2): 17-35
       Introduction: One of the core services the Winnipeg Regional Health Authority Virtual Library (WRHAVL) provides is literature searching for healthcare staff. Librarians send results to patrons as a reference list with links to the items identified. To improve access to the library's collection, and to understand patrons' engagement, the library developed a novel process for generating trackable links. Link and click data were analyzed to assess patron engagement with the literature search service.
    Methods: Leveraging the link management tool Short.io and a Visual Basic for Applications macro, the library uses an automated process to generate trackable links for literature search results sent to patrons. Click and link data are automatically collected in Short.io. Data from 19 October 2022 to 31 October 2024 was cleaned and exported into Excel for analysis. Logistic regression analysis was conducted in R.
    Results: Over the two-year study period, librarians sent 344 literature searches to patrons using this method. Patrons clicked at least one link in 194, or 56%, of reference lists. The percentage of total links clicked was higher for point of care and quick reference resources. Greater numbers of digital object identifier (DOI) and PubMed identifier (PMID) links included in a literature search predicted a decrease in the odds of someone clicking at least one link. Patrons clicked more than 20 links in only 10, or just 3%, of literature searches.
    Discussion: The data provides valuable insights into patron engagement with the WRHAVL's literature search service. The discovery that 44% of reference lists did not have a single click warrants further investigation.
    DOI:  https://doi.org/10.29173/jchla29914
  2. Database (Oxford). 2026 Jan 15. pii: baag034. [Epub ahead of print]2026
      Text-mining techniques are essential tools to efficiently access information within an increasing number of publications. The SIB Literature Services (SIBiLS) are dedicated to the text-mining and biocuration communities. SIBiLS features an extensive repository comprising ~40 million abstracts from MEDLINE, 8 million full-text articles accessible through the National Library of Medicine (NLM) via PMC, 29 million supplementary data files, and ~1 million taxonomic treatments sourced from Plazi and other contributors. The full-text articles are also containing original Open Access journals, which are not commonly indexed by PMC but relevant for research and scientific communities, the so-called PMC + collection. SIBiLS is thus a superset of the NLM collections. Several text-processing strategies have been implemented (e.g. tokenization, hyphenation, concept normalization), leading to the enrichment of SIBiLS with automatic annotations based on term mapping from ~30 ontologies. The aim is to improve literature triage and curation time by automating the extraction and organization of relevant information from large volumes of scientific texts. The annotation process has generated over 16 billion annotations. One of the goals of this enhancement is to improve the recall of rare contents, including genomic variants and information on rare diseases. It is also used in various contexts, including studying biotic interactions and curating genomic variants through dedicated front-end applications such as the BiotXplorer and Variomes. Annotations are performed using the BioC standard. By combining JATS and BioC, SIBiLS enables curators to perform high-precision evidence tracking at any textual representation levels via an original annotation schema (e.g. IDs, onto-terminological sources, provenance). An assessment of the annotation quality has been carried out for the different annotated entities. The average annotation accuracy is approaching 90% but exhibits high variation depending on both the entity and the source vocabulary.
    DOI:  https://doi.org/10.1093/database/baag034
  3. Explor Res Clin Soc Pharm. 2026 Dec;24 100837
      Prescription medication side effects are a prominent concern in healthcare, potentially compromising treatment outcomes and reducing patient quality of life. Increased exposure to side effect information, particularly from non-reputable sources, has been linked to increased side effect reports. An online survey involving a U.S. national adult sample that was census-benchmarked for gender, age, and race (N = 1313) was conducted to determine the sources patients use to obtain side effect information regarding prescription medications and their perceptions of these sources. Participants most frequently reported obtaining side effect information from healthcare providers, patient information leaflets, or the Internet, whereas pharmacists and interpersonal contacts were reportedly used less frequently. Health websites (e.g., WebMD or Mayo Clinic) were the most used online source for information; in contrast, social media, online forums, research articles, and government websites were less frequently consulted. Participants rated healthcare clinicians as the most trustworthy, reliable, accurate, and helpful source of side effect information, with more neutral attitudes expressed toward interpersonal and Internet sources. Additional analyses revealed significant effects of age, race, and gender on source use and perceptions, with women reporting greater engagement with nearly all sources compared to men, non-White participants more frequently relying on online forums and social media, and younger adults using healthcare practitioners and online forums more in comparison to older adults who favored leaflets and pharmacists. These findings highlight the variability in how individuals access and perceive side effect information.
    Keywords:  Health; Individual differences; Information seeking; Prescription medicine; Side effects
    DOI:  https://doi.org/10.1016/j.rcsop.2026.100837
  4. PLoS One. 2026 ;21(9): e0351281
      Dynamic Searchable Symmetric Encryption (DSSE) enables keyword search over encrypted data without revealing plaintext to the server. Arabic morphological richness - where a single root generates dozens of surface forms - creates substantial challenges for encrypted search: unprocessed vocabularies inflate encrypted-index size and transmission cost, and limit retrieval recall by failing to match morphological variants. This paper presents the first empirically grounded preprocessor selection framework for Arabic DSSE, derived from a systematic evaluation of four Arabic preprocessing strategies - the Khoja stemmer, ISRI stemmer, Lucene Arabic Analyzer, and Farasa segmenter - against a normalization-only baseline within the Incidence Matrix DSSE (IM-DSSE) scheme on the Khaleej Arabic news corpus. We evaluate vocabulary size, search latency, search quality, and retrieval breadth on an annotated 1,400-document corpus using 56 benchmark queries, and assess scalability on cloud infrastructure across corpus sizes up to 45,500 documents. Stemming reduces vocabulary by up to 83% relative to the normalization-only baseline, proportionally reducing encrypted-index size and transmission cost. We find that per-query search latency is driven primarily by result-set size rather than vocabulary size: the normalization-only baseline is fastest per query because it matches the fewest documents, while preprocessors that broaden retrieval incur higher latency. At scale, all five configurations - including the normalization-only baseline - operate successfully to 45,500 documents, and the result-set-size effect on latency persists. In terms of search quality, light stemming and morphological segmentation achieve the best trade-off between retrieval precision and recall, while root-based stemmers sacrifice precision without commensurate recall gains. Based on these findings, the framework provides actionable, evidence-based guidelines mapping deployment requirements - memory scalability, search latency, search quality, and retrieval breadth - to concrete preprocessor choices for Arabic DSSE system designers.
    DOI:  https://doi.org/10.1371/journal.pone.0351281
  5. Front Pediatr. 2026 ;14 1927624
       Background: Parents and caregivers increasingly use generative artificial intelligence chatbots for child health information. Pediatric vitamin D deficiency is a clinically relevant topic because advice about supplementation, testing, rickets, high-risk children, toxicity, and emergency symptoms can influence caregiver decisions.
    Objective: To compare the safety, medical accuracy, empathy, information reliability, educational quality, transparency, global quality, and readability of five large language model-driven chatbots when answering questions about pediatric vitamin D deficiency.
    Methods: This cross-sectional comparative study was reported with reference to CHART. An expert-curated 44-question set was selected from a 92-question candidate pool informed by search trends, caregiver-facing sources, clinical guidelines, and expert discussion. Each of the 44 questions was submitted once to each of five chatbot services-ChatGPT-5.5, Gemini 3.1 Pro, Qianwen 3.6-Plus, DeepSeek V4, and Doubao-Seed-2.0 Pro-using the same parent-oriented instruction. Responses were assessed for safety, accuracy, empathy, DISCERN, EQIP, JAMA criteria, GQS, and readability. Paired question-level differences were analyzed using Friedman tests, Kendall's W, and Cochran's Q; results were interpreted as a single-run, time-specific snapshot.
    Results: Inter-rater agreement was good to excellent (Fleiss' kappa = 0.842 for safety; ICCs 0.846-0.914 for other subjective metrics). Twenty of 220 responses (9.1%) were unsafe; safety did not differ significantly across models (Cochran's Q = 3.704, df = 4, P = 0.448). Significant inter-model differences were observed in accuracy, empathy, DISCERN, EQIP, JAMA, GQS, and all readability indices. ChatGPT had the highest median accuracy [5.00 (4.00, 5.00)] and DISCERN score [69.60 (65.25, 72.40)]; DeepSeek and Doubao had the highest empathy scores [5.00 (4.80, 5.00)]. ChatGPT and Doubao shared the highest EQIP median (86.00), Doubao had the highest GQS [5.00 (4.00, 5.00)], and Gemini had the highest FRES. JAMA scores were low across models.
    Conclusions: In this single-run evaluation, the five chatbot services showed domain-specific performance differences without a significant safety difference. Findings are exploratory rather than evidence of stable model superiority and support multidimensional evaluation with clinician involvement for high-risk or individualized pediatric advice.
    Keywords:  DISCERN; chatbot health advice; generative artificial intelligence; large language models; patient education; pediatric vitamin D deficiency; readability; rickets
    DOI:  https://doi.org/10.3389/fped.2026.1927624
  6. Front Psychiatry. 2026 ;17 1937506
       Background: Due to structural and socioeconomic barriers to accessing the healthcare system, individuals are increasingly turning to online platforms to obtain health-related information. With the advancement of artificial intelligence (AI), the use of AI for this purpose is also on the rise. It can also be said that seeking information from AI regarding psychotherapies is increasingly common.
    Objective: This study aimed to assess the informational quality and readability of AI-generated psychoeducational content about psychotherapy. To this end, we evaluated the informational quality and readability of LLM-generated Eye Movement Desensitization and Reprocessing (EMDR) psychoeducational content across different models.
    Methods: In this cross-sectional content analysis, EMDR-related search queries were identified through Google Trends, converted into patient-facing questions, and subsequently organized into four thematic categories, with five questions selected from each category for evaluation using ChatGPT (GPT-4o), Gemini (1.5 Pro), and DeepSeek (V3). Each question was entered in three separate sessions, yielding 180 responses. The informational quality was assessed by two blinded psychiatrists using the Ensuring Quality Information for Patients (EQIP) instrument. Readability was evaluated using the Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease (FRE) indices. Statistical analyses included linear mixed-effects models and ANOVA.
    Results: The overall mean EQIP score was 56.08 (SD = 4.43). Informational quality differed significantly across the models (p < 0.001), and no significant differences were observed across EMDR-related question categories. The highest informational quality scores were obtained by Gemini, followed by DeepSeek and ChatGPT. DeepSeek generated the most readable responses. Higher EQIP scores were seen to be associated with greater readability and longer response length. All the evaluated models produced responses that exceeded the commonly recommended readability levels for patient education materials. Overall inter-rater consistency was high for the total EQIP scores.
    Conclusion: Although all the LLMs evaluated were able to generate psychoeducational information about EMDR, substantial differences were observed between the models in respect of informational quality and readability. The findings suggest that psychoeducational quality depends more strongly on the language model used than on the specific EMDR-related topic being addressed. LLMs may serve as supplementary psychoeducational resources but should not be considered as substitutes for professional psychotherapy or individualized clinical care.
    Keywords:  EMDR therapy; artificial intelligence; digital mental health; informational quality; large language models; psychoeducation; readability
    DOI:  https://doi.org/10.3389/fpsyt.2026.1937506
  7. Nutr Hosp. 2026 Sep 16.
       BACKGROUND: large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality, and readability.
    METHODS: in this cross-sectional comparative study, 50 public facing questions about nutrition in cirrhosis were submitted once to each of four LLMs (ChatGPT 5.5, Claude Opus 4.7, DeepSeek V4 Pro, and Gemini 3.1 Pro Thinking) between May 10 and May 16, 2026. Three raters independently assessed 200 responses for safety, accuracy, empathy, and information quality. Readability was assessed with six indices.
    RESULTS: a total of 200 responses were evaluated. Safety coding showed high inter-rater agreement (Fleiss' kappa = 0.864; 95 % CI, 0.755-0.958). ICC values for other manually scored outcomes ranged from 0.860 to 0.892. Potentially unsafe responses occurred in all models, ranging from 6.0 % to 12.0 %, with no significant difference across models (raw p = 0.599). Accuracy differed significantly across models (raw p = 0.003), with the highest median score for ChatGPT 5.5. Empathy also differed significantly (raw p < 0.001), with higher median scores for Claude Opus 4.7 and Gemini 3.1 Pro Thinking. DISCERN, EQIP, JAMA and all readability indices differed significantly across models, while GQS did not.
    CONCLUSION: current LLMs can provide useful general information about nutrition in cirrhosis, but none performed consistently well across safety, transparency, and readability. Their responses should supplement, not replace, guidance from hepatology and nutrition professionals.
    DOI:  https://doi.org/10.20960/nh.06997
  8. Comput Biol Med. 2026 Sep 25. pii: S0010-4825(26)00513-5. [Epub ahead of print]216 111947
       OBJECTIVE: Wilson's disease is a rare genetic disorder with a large but highly fragmented body of biomedical literature spanning more than a century of research. Although thousands of articles are available through PubMed, this knowledge remains difficult to access in a consolidated, query-driven manner for researchers, clinicians, and students. Existing large language models (LLMs) offer conversational access to information but often lack domain grounding, leading to hallucinations and unreliable responses in biomedical settings. In this work, we introduce WilsonLitQA, a literature-grounded retrieval-augmented question answering resource that enables the scientific community to query the complete PubMed literature on Wilson's disease, capturing canonical biological, clinical, and pathophysiological knowledge related to the disease.
    METHODS: Our framework integrates hybrid information retrieval with generative language models, enabling users to query the Wilson's disease literature as a unified knowledge resource rather than as isolated papers. We have systematically evaluated multiple open-source and proprietary LLMs under zero-shot and few-shot in-context learning settings, assessing their ability to deliver accurate, complete, and non-hallucinated responses when supported by retrieval.
    RESULTS: Our results demonstrate that retrieval grounding substantially improves biological accuracy and clinical relevance, and that smaller, deployable language models (on the order of 7B parameters) can perform competitively when paired with a well-designed retrieval pipeline.
    CONCLUSION: To support transparency, reproducibility, and community-driven development, the complete WilsonLitQA implementation is publicly available (https://dhanjal-lab.iiitd.edu.in/wilsonlitqa.html). Together, this work positions retrieval-augmented question answering as a consolidated, continuously extensible knowledge interface for rare diseases, offering the scientific community a practical tool to navigate, interpret, and query the growing biomedical literature on Wilson's disease.
    Keywords:  Biomedical literature; Large language models; Question answering; Retrieval; Wilson's disease; WilsonLitQA
    DOI:  https://doi.org/10.1016/j.compbiomed.2026.111947
  9. Healthcare (Basel). 2026 Sep 21. pii: 3119. [Epub ahead of print]14(18):
      Background/Objectives: Patients use artificial intelligence (AI) chatbots for medical information, but response quality, source transparency, and readability for primary aldosteronism (PA) remain uncertain. This study compared four chatbot configurations and assessed whether their responses met patient-education reading levels. Methods: During a single-time-point snapshot on 27 July 2026, 17 questions derived from Medical Subject Headings and five-year worldwide Google Trends data were submitted once to GPT-5.5 through ChatGPT Plus, Copilot in Smart mode, Gemini 3.5 Flash, and the default Perplexity model, yielding 68 responses. DISCERN, the Ensuring Quality Information for Patients (EQIP) instrument, the Journal of the American Medical Association (JAMA) benchmarks, and the Global Quality Score (GQS) assessed structural information quality; six formulas assessed readability. These instruments did not assess statement-level factual accuracy. Differences were examined using Friedman tests and paired Wilcoxon signed-rank tests with Holm adjustment. Results: Scores differed for DISCERN (χ2 = 16.717, p < 0.001), EQIP (χ2 = 31.125, p < 0.001), JAMA (χ2 = 48.851, p < 0.001), and GQS (χ2 = 9.874, p = 0.020). Copilot had the highest median DISCERN and EQIP scores (43.00 and 75.00), followed by Perplexity (42.00 and 70.00). Median JAMA scores were 0 for ChatGPT and Gemini and 1 for Copilot and Perplexity. These instrument-based differences do not establish greater factual accuracy, guideline concordance, or clinical superiority; moreover, the full-set DISCERN comparison has limited interpretability because only two questions were explicitly treatment-related. Item-level kappa values ranged from 0.844 to 0.924, and total-score ICC(2,1) values ranged from 0.846 to 0.883. No readability measure met the sixth-grade benchmark. Conclusions: In these configurations, chatbots produced PA information. However, source transparency remained low, and language was overly complex. Relative score differences should not be interpreted as factual superiority. Patient-facing use requires verified sources, scope limits, human oversight, and routes to care.
    Keywords:  artificial intelligence chatbot; health information quality; large language model; patient education; primary aldosteronism; readability
    DOI:  https://doi.org/10.3390/healthcare14183119
  10. Cancers (Basel). 2026 Sep 20. pii: 3047. [Epub ahead of print]18(18):
       BACKGROUND: At present, there are very limited studies evaluating autonomous artificial intelligence (AI) agents in oral oncology or healthcare education, and no studies have directly compared traditional frontier large language models (LLMs) with autonomous AI agents in the field of pediatric oral oncologic pathology. This study evaluated the performance of AI LLMs and an autonomous AI scientific agent in generating engaging and accessible patient information sheets for common pediatric oral pathologic conditions and rare head and neck tumors. The development of high-quality patient education materials is particularly important in pediatric pathology because parents must often navigate complex and emotionally sensitive diagnoses, including rare tumors and developmental lesions, for which accessible, patient-friendly educational resources are frequently unavailable. AI-generated patient information sheets therefore represent a potential strategy to improve communication, understanding, and shared decision-making for families facing these uncommon conditions.
    METHODS: AI-generated patient information sheets from popular frontier chatbots on various pediatric pathological conditions and rare tumors were evaluated by five platforms (Eli5a v2.0, ChatGPT-5.4, Claude Sonnet 4.6, Perplexity and Doximity), in a blinded fashion, by three expert evaluators using the Global Quality Score (GQS), DISCERN score, understandability score and actionability score using PEMAT, and the Flesch-Kincaid Grade Level. Because the same 20 conditions were assessed on every platform, observations were paired; platforms were compared using Friedman tests with paired Wilcoxon signed-rank post hoc tests and Holm correction, and interrater reliability was quantified using an absolute-agreement intraclass correlation coefficient (ICC).
    RESULTS: Platform differences were significant for all five outcomes (all p < 0.001). Claude Sonnet 4.6 obtained the best overall information quality (GQS 4.25 ± 0.39), scoring significantly higher than all other platforms, while Perplexity, Eli5a v2.0 and ChatGPT-5.4 performed similarly to one another and better than Doximity. For information reliability, Doximity (DISCERN 70.50 ± 3.20) and Claude Sonnet 4.6 (70.25 ± 3.02) performed best, while Eli5a v2.0 recorded the lowest DISCERN score (54.50 ± 4.26). Eli5a v2.0 obtained the best results for patient-centered communication, with the highest understandability (PEMAT-U 95.05 ± 1.23) and readability (FKGL 5.45 ± 0.51) and an actionability score among the highest of the five platforms (PEMAT-A 89.00 ± 2.62, not significantly different from Perplexity or Doximity). Interrater absolute agreement for GQS was poor to moderate (ICC(2,1) = 0.236; ICC(2,3) = 0.481).
    CONCLUSION: No single platform was superior across all evaluated domains. Claude Sonnet 4.6 led in overall quality whereas Eli5a v2.0 achieved the highest understandability and the most appropriatereading grade level. Eli5a's actionability was high but did not differ significantly from Perplexity or Doximity. These findings indicate a trade-off between information quality and reliability on one hand and lay accessibility on the other. Platform selection should therefore be matched to the communication task, and all AI-generated patient materials require clinician review before use.
    Keywords:  artificial intelligence; autonomous; bot; engagement; health information; large language models
    DOI:  https://doi.org/10.3390/cancers18183047
  11. Mult Scler Relat Disord. 2026 Sep 23. pii: S2211-0348(26)00972-7. [Epub ahead of print]115 107937
       OBJECTIVE: This study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise.
    METHOD: A total of 75 questions were evaluated. Expert physiotherapists analysed the responses utilizing the Global Quality Score (GQS), Modified DISCERN (mDISCERN), Likert Accuracy Scale, and the Flesch Reading Ease (FRE).
    RESULTS: While 61.3% of ChatGPT responses were classified as high quality, this rate was 24% for Gemini (p < 0.001). ChatGPT-generated responses demonstrated significantly higher overall accuracy and mDISCERN scores than those generated by Gemini (p < 0.001). Categorical analyses revealed that the responses generated by ChatGPT received higher accuracy in the domains of exercise planning, safety and risk management, as well as participation and self-management. Additionally, ChatGPT-generated responses received significantly higher mDISCERN scores in the exercise planning and safety categories. No significant difference was observed between the two models regarding overall FRE scores (p > 0.05). The mean FRE scores for ChatGPT and Gemini were 41.70 and 38.21, respectively, with responses from both models classified at a "difficult" readability level.
    CONCLUSION: Within the scope of this study, ChatGPT-generated responses demonstrated higher quality, accuracy, and reliability than those generated by Gemini. Nonetheless, patient accessibility could be limited by the poor readability metrics observed in both tools. While AI-based chatbots may show promise in reinforcing patient education for MS populations, they must not substitute specialized medical professionals during clinical decision-making.
    Keywords:  Artificial intelligence; ChatGPT; Exercise; Google Gemini; Multiple sclerosis
    DOI:  https://doi.org/10.1016/j.msard.2026.107937
  12. Medicine (Baltimore). 2026 Sep 25. 105(39): e50839
      Flatfoot (pes planus) is common in children and adults. Although often asymptomatic, it may alter lower-limb biomechanics and contribute to pain or injury. Patients and caregivers increasingly use artificial intelligence chatbots such as chat generative pretrained transformer (ChatGPT) and Google Gemini for medical information, yet their responses on flatfoot remain insufficiently studied. This study compared the factual accuracy, added value, omissions, and readability of responses generated by ChatGPT and Google Gemini to standardized patient-oriented questions. In this cross-sectional comparative study, 15 standardized patient-oriented questions were developed from clinical guidelines, peer-reviewed literature, and educational resources from established orthopedic societies, then refined by 2 orthopedic surgeons. Each question was submitted to ChatGPT (OpenAI) and Google Gemini (Google LLC) in July 2025 using identical wording in independent chat sessions. Responses were anonymized and independently rated by 2 board-certified orthopedic surgeons for factual accuracy, added value, and omissions using an investigator-developed 5-point rubric. Discrepant assessments were reviewed by a 3rd senior orthopedic surgeon. Primary analyses used mean scores of the 2 reviewers. Readability was assessed using the Flesch-Kincaid grade level and Flesch reading ease score. Paired differences were analyzed using the Wilcoxon signed-rank test, with P < .05 considered significant. Both models produced generally accurate responses, with mean factual accuracy scores above 4/5. Gemini scored higher than ChatGPT for factual accuracy (4.7 ± 0.2 vs 4.2 ± 0.3, P = .01), added value (4.6 ± 0.3 vs 4.0 ± 0.4, P = .02), and omissions (4.8 ± 0.2 vs 3.9 ± 0.4, P < .001), with higher scores indicating fewer clinically relevant omissions. Gemini also generated longer responses (165 ± 20 vs 120 ± 15 words, P < .001) and showed better readability, with lower Flesch-Kincaid grade level (9.2 ± 1.1 vs 11.8 ± 1.4, P < .001) and higher Flesch reading ease score (59.4 ± 6.2 vs 42.5 ± 5.3, P < .001). Both models generated generally accurate answers to standardized flatfoot questions. Gemini performed better across evaluation domains and readability measures. As these findings are time-specific, repeated assessments and direct patient and caregiver evaluations are needed before broader patient-facing use can be recommended.
    Keywords:  ChatGPT; Google Gemini; accuracy; artificial intelligence; flatfoot; orthopedics; patient education; pes planus; readability
    DOI:  https://doi.org/10.1097/MD.0000000000050839
  13. PLoS One. 2026 ;21(9): e0359170
      Pediatric chest pain represents a frequently encountered symptom in emergency departments and outpatient clinics throughout childhood, often generating substantial parental anxiety. This study aims to comparatively analyze the readability, information quality, and scientific reliability of textual contents generated by the artificial intelligence (AI) chatbots ChatGPT, Gemini, and Perplexity regarding pediatric chest pain, utilizing multidimensional analytical indices. Out of the top 25 queries with the highest global search volume on Google Trends, 17 unique keywords meeting the predefined inclusion criteria were filtered on June 1, 2026. These inquiries were directed to all three AI platforms in distinct, independent user sessions, yielding a total of 51 textual responses. Linguistic readability levels were calculated across six separate digital interfaces using the FKGL, FRES and other formulas, and the results were benchmarked against the sixth-grade reading level" threshold. Scientific reliability was audited via the Modified DISCERN and JAMA benchmarks, while content quality was examined using the GQS and EQIP instruments. The median readability scores computed across all three AI models were found to be statistically and significantly above the targeted sixth-grade comprehension threshold, indicating a high level of difficulty (p < 0.001). Post-hoc pairwise evaluations revealed that ChatGPT and Gemini offered linguistically more accessible and structurally less complex textual architectures; conversely, Perplexity exhibited a significantly more difficult and academic linguistic framework (p < 0.0167). On the other hand, regarding content quality and source credibility, Perplexity demonstrated remarkably superior and more optimized scores across all evaluated instruments, namely GQS (p = 0.002, p = 0.001), JAMA (p < 0.001, p < 0.001), mDISCERN (p = 0.005, p = 0.001), and EQIP (p = 0.003, p < 0.001), compared directly to both ChatGPT and Gemini, respectively. Although generative AI formulations harbor considerable potential to provide extensive data on pediatric chest pain, the sophisticated, university-level linguistic architecture of these texts poses a substantial access barrier for individuals who lack proficient digital health literacy skills. While Perplexity achieved the performance closest to the "gold standard" in terms of informational accuracy, no model is currently mature enough to substitute for a professional medical consultation due to observed citation biases and omissions. It is of paramount importance that clinicians actively guide families to scrutinize online health data through a critical lens.
    DOI:  https://doi.org/10.1371/journal.pone.0359170
  14. Front Public Health. 2026 ;14 1914167
       Objective: Myocardial infarction (MI) is an acute, life-threatening cardiovascular disease, and high-quality, accessible public health education is vital for emergency management. This study systematically evaluates the quality, transparency, clinical accuracy, patient safety, and readability of information generated by different large language models (LLMs) in responding to MI-related public inquiries.
    Methods: Twenty-five representative MI patient education questions were submitted to Gemini 3.5 Flash, Claude Opus 4.8, and ChatGPT 5.5. The generated information was independently evaluated by two cardiologists using four validated tools (DISCERN, EQIP, GQS, and JAMA) alongside a strict clinical safety assessment. Text readability was concurrently assessed utilizing six established metrics (FRES, ARI, GFI, CLI, FKGL, and SMOG).
    Results: Significant variations were observed in the quality, transparency and readability of information generated by the evaluated LLMs. Regarding quality and transparency, significant overall differences were noted among models in DISCERN (p < 0.001) and EQIP (p < 0.001) scores, whereas no significant differences were found in GQS and JAMA benchmarks. For DISCERN, Claude achieved significantly higher scores than both Gemini and ChatGPT. In the EQIP assessment of completeness and clarity, Claude and Gemini scored significantly higher than ChatGPT, though all models attained a "good" rating. Additionally, all models exhibited poor performance on the JAMA benchmark, indicating critical deficits in information transparency. Notably, LLMs sometimes generated incomplete, incorrect, or even potentially harmful information. Regarding readability, although Claude generated relatively more comprehensible text, all models failed to meet the recommended sixth-grade reading benchmark, indicating high reading difficulty.
    Conclusion: While LLMs can generate structurally clear and logically coherent foundational content for MI-related queries, they occasionally produce clinically inappropriate directives. Furthermore, the texts generated by these models are overly complex, creating substantial reading barriers for the general public. Consequently, under zero-shot and English-language testing conditions, the current LLMs are not yet capable as standalone health education tools for MI.
    Keywords:  information quality; large language models; myocardial infarction; public health education; readability
    DOI:  https://doi.org/10.3389/fpubh.2026.1914167
  15. Work. 2026 Sep 23. 10519815261488805
      BackgroundPectus excavatum and carinatum are the most common anterior chest wall deformities in children and adolescents. With growing awareness and easy access to online resources, patients increasingly seek information about these conditions from artificial intelligence platforms.ObjectiveThe aim of this study is to evaluate the artificial intelligence model Chat Generative Pre-trained Transformer (ChatGPT) by analyzing its responses to frequently asked questions about pectus deformities and assessing their accuracy and comprehensiveness.MethodsWebsites frequently used by patients with pectus and their families, as well as quality-of-life questionnaires, have been examined by the authors, and a total of 99 questions have been collected. ChatGPT's responses were evaluated by a physiatrist and a pediatric surgeon using a 4-point grading system: 1) correct and comprehensive, 2) correct but not comprehensive, 3) partially correct and partially incorrect, 4) completely incorrect.Results51 questions (51.5%) received correct and comprehensive answers, 40 questions (40.4%) received correct but not comprehensive answers, 8 questions (8.1%) received partially incorrect answers, and none received completely incorrect answers. ChatGPT was least successful in the pectus carinatum treatment and follow-up category, answering 19 questions with 21.1% correct, but not comprehensively.ConclusionsChatGPT's responses about pectus deformities have been found highly accurate, and more than half of the responses are comprehensive. It can be a useful and fast source of information for basic patient education. However, every patient is unique, and treatment and follow-up require a holistic evaluation. This point seems to be a limitation of ChatGPT and underscores the indispensable role of specialized medical consultants.
    Keywords:  consumer health information; funnel chest; generative artificial intelligence; large language models; pectus carinatum; thoracic wall
    DOI:  https://doi.org/10.1177/10519815261488805
  16. R I Med J (2013). 2026 Oct 01. 109(10): 35-40
       OBJECTIVES: Chatbots have been increasingly recognized as modern tools that provide patients with reliable health-related information. This study aimed to evaluate and compare ChatGPT-3.5 and GPT-4's ability to answer glenohumeral osteoarthritis-related questions.
    METHODS: Fifteen questions were derived from the 2020 AAOS Clinical Practice Guidelines for the Surgical Management of Glenohumeral Joint Osteoarthritis. Questions were categorized into three groups: risk factors, implant/intraoperative considerations, and pain/functional outcomes. ChatGPT-3.5 and GPT-4 were prompted with these questions, and responses were evaluated by four fellowship-trained shoulder and elbow surgeons. Each response was rated on a scale (scores:1-5) based on relevance, accuracy, clarity, completeness, and evidence-based support. Data was analyzed descriptively and statistically to compare the scores between ChatGPT-3.5 and GPT-4.
    RESULTS: Average score for ChatGPT-3.5 was 19.7/25, with "Risk Factor" prompts achieving the highest mean score. GPT-4 averaged 18.7/25, with "Functional Outcomes" prompts scoring highest. However, there were no statistically significant differences between different prompt themes for GPT-3.5 and GPT-4. "Clarity" category received the highest score for GPT-3.5, while "Relevance" was highest for GPT-4. Both models scored lowest on "Evidence-based" prompts. On the Flesch- Kincaid scale, GPT-3.5 responses had a significantly higher score of 18.3 compared to GPT-4's 15.4, indicating a more difficult reading level in GPT-3.5's responses.
    CONCLUSION: Both ChatGPT-3.5 and GPT-4 performed adequately in providing well-informed medical responses to patient queries about glenohumeral osteoarthritis. Future chatbot versions should focus on providing evidence-based content through systematic and reliable reviews of literature, in an accessible readable manner.
    Keywords:  Flesch-Kincaid; artificial intelligence; chatbot; health information; patients
  17. J Pharm Technol. 2026 Sep 18. 87551225261487336
       Background: The National Institutes of Health (NIH) recommends patient education materials (PEMs) be written below a seventh-grade reading level. Appropriate reading levels are important, as patients with mental health conditions and low health literacy are at greater risk of poor health outcomes.
    Objective: To assess and compare the readability level of NIH mental health PEMs using the Readable website and 2 artificial intelligence (AI) chatbots.
    Methods: Readability metrics (Flesch-Kinkaid reading level, Gunning Fog, Coleman-Liau, SMOG, and Automated Readability Index) of NIH mental health PEMs were evaluated using Readable, ChatGPT, and Claude AI in this cross-sectional descriptive study. Common readability issues were collected. Coloring pages and activity worksheets for children were excluded.
    Results: Forty PEMs were included in the final analysis. The average Flesch-Kincaid reading levels according to Readable, ChatGPT, and Claude were 8.74 (SD 1.85), 10.44 (SD 3.13), and 9.58 (SD 2.49), respectively. Average Gunning Fog, Coleman-Liau, SMOG, and Automated Readability Index scores provided by Readable were 10.51 (SD 2.09), 14.29 (SD 2.36), 11.27 (SD 1.46), and 10.18 (SD 2.13), respectively. The most common readability issue was sentences exceeding 20 syllables. Limitations included a single-time-point readability assessment of NIH mental health PEMs limited to 2 AI chatbots.
    Conclusion: NIH mental health PEMs consistently exceeded the seventh-grade reading level and should be revised to adhere to NIH recommendations. Readability metrics varied across Readable and AI chatbot platforms. Improving PEM readability enhances patient understanding, promotes informed decision-making, and reduces disparities linked to limited health literacy, particularly within mental health populations.
    Keywords:  artificial intelligence; chatbot; mental health; patient education; readability
    DOI:  https://doi.org/10.1177/87551225261487336
  18. Cureus. 2026 Aug;18(8): e115019
       BACKGROUND: Nail-patella syndrome is a rare autosomal dominant multisystem disorder with musculoskeletal, renal, and ocular manifestations. Patients and families may rely on online resources, but the completeness and accuracy of this information have not been well characterized. This study evaluated English-language online information about nail-patella syndrome.
    METHODS: A structured cross-sectional pilot search conducted on August 6, 2026, identified 24 candidate domains. Fifteen independent patient-facing websites were included in the primary analysis. Pages were grouped by source type and assessed using a 15-item Nail-Patella Syndrome Specific Content Score (range 0-15) and an accuracy-completeness score in which each item was rated 0 (absent), 1 (present but incomplete or partly inaccurate), or 2 (present and broadly accurate; range 0-30). GeneReviews and peer-reviewed literature formed the reference standard.
    RESULTS: The mean presence score was 13.2 of 15 (standard deviation 1.8; range 10-15), and the mean accuracy-completeness score was 23.7 of 30 (standard deviation 5.3; range 14-30). Government/non-profit websites had the highest mean accuracy-completeness score (26.0), followed by healthcare-provider (25.0), media/general-information (23.7), and commercial websites (20.0); differences were not statistically significant (Kruskal-Wallis H(3)=2.58, p=0.462). Nail, knee, and elbow manifestations were consistently described. Epidemiology was present on eight websites, while diagnosis/investigations, genetic counselling, and surveillance were each present on 11. Renal manifestations were mentioned on all websites, but only 10 were fully accurate; glaucoma was omitted by one website.
    CONCLUSIONS: This pilot study identified specific gaps in online patient-facing information about nail-patella syndrome, particularly concerning diagnosis, renal and ocular surveillance, and balanced management. These findings highlight areas for improvement in patient resources rather than establishing a definitive ranking of individual websites.
    Keywords:  glaucoma; infodemiology; information accuracy; lmx1b; nail-patella syndrome; nephropathy; online health information; patient education
    DOI:  https://doi.org/10.7759/cureus.115019
  19. PeerJ. 2026 ;14 e21730
       Background: YouTube is one of the most widely accessed digital platforms for health-related information, particularly among children exposed to animated content. However, the educational quality of cartoon-based videos on children's tooth brushing remains uncertain. This study evaluated the educational value and content characteristics of YouTube cartoon videos related to children's tooth brushing.
    Methods: A cross-sectional content analysis was conducted on the first 200 videos retrieved on December 28, 2025, using the keywords "children," "tooth brushing," and "cartoon." After applying inclusion and exclusion criteria, 109 videos were included. Educational quality was assessed using previously published criteria and categorized into four levels. Inter-rater reliability was calculated using the intraclass correlation coefficient (ICC). Associations between educational score and video characteristics were examined using Spearman correlation and ordinal logistic regression.
    Results: Most videos were classified as slightly useful, and only one video (0.9%) was categorized as highly useful. The mean educational score was 2.04 ± 1.37. Tooth brushing demonstrations and daily brushing frequency were commonly presented; however, essential preventive topics such as appropriate toothpaste amount, parental supervision, floss use, and recommended starting age were frequently absent. Inter-rater reliability was good (ICC = 0.76). A weak but significant negative correlation was observed between video duration and educational score (r =  - 0.195, p = 0.042). Ordinal logistic regression showed that longer video duration was independently associated with lower odds of higher educational quality (OR = 0.33, 95% CI [0.11-0.99], p = 0.048). No significant associations were found between educational score and views, likes, interaction index, or upload duration. Cartoon-based YouTube videos on children's tooth brushing demonstrate limited educational quality despite high accessibility. Popularity metrics do not reflect educational value, underscoring the need for evidence-based, structured digital oral health content for children.
    Keywords:   Cartoon videos; Children; Content analysis; Oral hygiene; YouTube
    DOI:  https://doi.org/10.7717/peerj.21730
  20. PLoS One. 2026 ;21(9): e0357587
       BACKGROUND: Myopia is a public health concern, and its prevalence is estimated to increase in the near future. Multiple myopia control strategies have been developed to slow myopia progression and reduce potential ocular alterations. Public interest in myopia and its management has increased, with parents seeking information on myopia from available sources, such as the internet. The goal of this study was to systematically evaluate the quality, reliability, and educational value of YouTube videos for myopia control in children and their alignment with current scientific evidence and clinical guidelines.
    METHODS: A comprehensive search was conducted on YouTube in February 2025 using four primary keywords related to myopia control in children, both in English and Spanish. The first 50 videos for each keyword were screened; of the resulting 400 videos, 235 met the inclusion criteria. Three optometrists independently assessed the videos via the modified DISCERN (mDISCERN), the Journal of the American Medical Association (JAMA) benchmarks, and the Global Quality Score (GQS). Interobserver reliability was evaluated using two-way mixed-effects intraclass correlation coefficients (ICC) with absolute agreement. Quantitative engagement metrics (views, likes, dislikes, the view ratio, and the video power index (VPI) were also recorded. Non-parametric tests were used for subgroup comparisons by language, topic, source, and presenter gender, with Bonferroni adjustments applied for multiple comparisons.
    RESULTS: Among the 235 videos analysed, 52.8% were in English, and 47.2% were in Spanish. Overall video quality and reliability were highly variable. Interobserver reliability across the entire dataset was good for GQS (ICC = 0.832) and mDISCERN (ICC = 0.750), but poor for JAMA (ICC = 0.447). Good interevaluator agreement was maintained across all tools for Spanish videos (ICC range: 0.701-0.803), whereas English videos showed lower agreement for mDISCERN (ICC = 0.768) and JAMA (ICC = 0.536), but excellent agreement for GQS (ICC = 0.849). Spanish-language videos scored significantly higher on mDISCERN (P = 0.026), while JAMA and GQS scores showed no significant differences between languages after adjusting for multiple comparisons. Videos from universities or professional organisations were of higher quality and reliability but constituted a minority of the sample. The presenter's gender did not impact quality or engagement. English-language videos had higher view ratios and VPI values, indicating greater visibility, although not necessarily higher scientific quality.
    CONCLUSIONS: YouTube provides a wide range of content on pediatric myopia control, but its overall quality, realiability and scientific alignment are inconsistent. The use of standardized evaluation tools reveals significant differences in subjective assessment depending on the scoring instrument and video language. The increased participation of healthcare professionals in creating evidence-based videos is needed to increase public health literacy.
    DOI:  https://doi.org/10.1371/journal.pone.0357587
  21. Medicine (Baltimore). 2026 Sep 25. 105(39): e50914
      Lumbar spinal stenosis (LSS) is a common degenerative condition that impairs mobility and quality of life, particularly in older adults. Video-sharing platforms are increasingly used for health information, yet the quality and reliability of LSS-related videos and cross-platform differences remain unclear. This study compares the content, uploader characteristics, and information quality of LSS-related videos on YouTube, TikTok, and Bilibili and examines associations between video quality and audience engagement. A cross-sectional search of YouTube, TikTok, and Bilibili was conducted on February 1, 2026, using relevant Chinese and English keywords. Of the retrieved videos, 253 met the eligibility criteria. Reliability, overall quality, understandability, actionability, and audiovisual quality were assessed using the modified DISCERN instrument, Global Quality Score, Patient Education Materials Assessment Tool understandability and Patient Education Materials Assessment Tool actionability score scores, and Video Information and Quality Index (VIQI). The sample comprised 89 YouTube, 80 Bilibili, and 84 TikTok videos, with median durations of 389, 171, and 103 seconds, respectively. TikTok videos were the shortest and showed relatively high engagement across available metrics. YouTube videos more frequently covered etiology and prevention and were predominantly uploaded by hospitals/departments and verified physicians; Bilibili and TikTok videos focused mainly on treatment, and TikTok had the highest proportion of verified physician uploaders. YouTube videos had higher modified DISCERN instrument and Global Quality Scores. Total VIQI scores were higher on YouTube and TikTok than on Bilibili, whereas Patient Education Materials Assessment Tool actionability scores did not differ significantly across platforms. Associations between video quality and audience engagement varied by platform, with VIQI showing the most consistent positive correlations. The quality and dissemination characteristics of LSS-related videos differed across platforms. YouTube videos were more reliable, TikTok videos showed greater dissemination appeal, and Bilibili showed no clear advantage in most quality metrics. High audience engagement did not necessarily indicate medical reliability. Platform recommendation systems should consider uploader qualifications, content reliability, and presentation quality to promote high-quality LSS information.
    Keywords:  Bilibili; TikTok; lumbar spinal stenosis; patient education; social media; video quality
    DOI:  https://doi.org/10.1097/MD.0000000000050914
  22. Front Digit Health. 2026 ;8 1956383
       Background: Myelodysplastic syndromes (MDS) are common hematologic malignancies among older adults. With the rapid growth of short-video platforms, patients increasingly seek MDS-related health information online; however, the quality of MDS-related health information available on short-video platforms has not been systematically evaluated.
    Objective: This study assessed the quality and reliability of MDS-related short videos across three major Chinese platforms and examined the relationship between content quality and user engagement.
    Methods: Between June 25 and June 30, 2026, a total of 210 MDS-related videos were systematically collected from TikTok (Chinese version: Douyin), Bilibili, and Xiaohongshu (70 videos per platform). Video metadata were extracted, and content quality was independently evaluated using three validated instruments: the modified DISCERN (mDISCERN), the Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Scale (GQS). An overall quality score (SUM) was calculated by combining the three assessment tools. Spearman correlation analysis was performed to examine the associations between user engagement metrics and content quality.
    Results: Bilibili videos demonstrated the highest overall educational quality (mean SUM score: 9.79 ± 2.35), whereas TikTok generated the greatest user engagement (median number of likes: 81.00). Videos produced by professional creators achieved significantly higher scores across all quality assessment instruments than those produced by non-professional creators (mean SUM score: 8.81 vs. 5.03, p < 0.001). No significant associations were observed between any user engagement metric and content quality. Video duration was the only variable positively correlated with content quality (r = 0.332, p < 0.01).
    Conclusions: User engagement was disconnected from information quality in MDS-related short videos. Professional creators provided greater educational value without consistently receiving commensurate engagement. Digital health communication should prioritize expert-led content, transparent source attribution, and quality-informed dissemination.
    Keywords:  digital health literacy; myelodysplastic syndromes; quality assessment; short-video platforms; social media platforms
    DOI:  https://doi.org/10.3389/fdgth.2026.1956383
  23. Medicine (Baltimore). 2026 Sep 25. 105(39): e50773
      Rotator cuff injury (RCI) is a common age-related musculoskeletal disorder, yet public awareness remains limited. With short-video platforms emerging as key health information sources, concerns about content quality have grown. This study evaluated the quality, reliability, and comprehensiveness of RCI-related videos on TikTok and Bilibili and examined the impact of uploader characteristics. On October 13, 2025, Chinese-language videos containing "rotator cuff injury" were retrieved from TikTok and Bilibili. Following strict inclusion criteria, 231 videos were analyzed. Video characteristics, engagement metrics (likes, comments, shares, collections), and uploader information were collected. Quality and content integrity were assessed using the Video Information and Quality Index (VIQI), Global Quality Scale (GQS), and modified DISCERN (mDISCERN). Statistical comparisons and Spearman correlation analyses were performed. Of 231 videos, 96 were from Bilibili and 135 from TikTok. TikTok videos were shorter (median 93 seconds vs 234.5 seconds, P < .001) but demonstrated higher engagement across all metrics (P < .001). Physicians contributed most content (TikTok: 79%, Bilibili: 64%). Diagnosis and treatment dominated, while epidemiology and prevention were rarely addressed. GQS and mDISCERN scores did not differ between platforms. TikTok scored higher on VIQI information flow (P < .001) and visual quality (P = .044); Bilibili excelled in information accuracy (P = .006). Individual user videos showed slightly higher GQS (P < .05) and accuracy (P = .033); physician videos received more comments (P = .002) and shares (P = .042). VIQI and GQS correlated positively with engagement metrics (VIQI: R = 0.21-0.40; GQS: r = 0.20-0.28; all P < .001), whereas mDISCERN showed weak associations. Video duration correlated positively with quality scores (GQS: r = 0.56; VIQI: R = 0.34; mDISCERN: r = 0.33; all P < .001) but negligibly with engagement. The analysis of RCI-related videos on TikTok and Bilibili reveals moderate content quality and a significant lack of preventive information. These findings underscore the need for strategies to enhance the scientific rigor and comprehensiveness of health information on short-video platforms.
    Keywords:  Bilibili; TikTok; health communication; rotator cuff injury; short video
    DOI:  https://doi.org/10.1097/MD.0000000000050773
  24. Medicine (Baltimore). 2026 Sep 25. 105(39): e50873
      Flap transplantation is a crucial reconstructive surgical technique, yet public understanding of its principles and postoperative management remains limited. With the rapid rise of short-form video platforms, social media has become an important medium for disseminating medical information. This study aimed to evaluate the content characteristics, educational quality, and reliability of flap transplantation surgery videos on TikTok and Bilibili. A cross-sectional content analysis was conducted to evaluate short videos related to flap transplantation surgery on TikTok and Bilibili. Videos were retrieved on September 25, 2025, using the Chinese keyword "" ("flap transplantation surgery"), and searches were performed on the same day to reduce algorithm-related bias. After excluding duplicate, irrelevant, and commercial content, 236 videos were included. Two independent researchers extracted and verified data on video characteristics and engagement metrics. Educational content and information quality were assessed using the Global Quality Scale (GQS), modified DISCERN (mDISCERN), and Journal of the American Medical Association (JAMA) criteria. Statistical analyses were performed using Statistical Package for the Social Sciences version 27.0 with the Mann-Whitney U test, Kruskal-Wallis test, and Spearman correlation. Statistical significance was set at P < .05. Overall, the videos demonstrated moderate quality, with median scores of 3.00 (interquartile range 3.00-4.00) for GQS, 3.00 (2.00-3.00) for mDISCERN, and 3.00 (3.00-4.00) for JAMA. TikTok videos were significantly shorter (52.50 seconds vs 162.00 seconds, P < .001) but achieved higher GQS, mDISCERN, and JAMA scores (P < .001) compared with Bilibili. Specialist physicians accounted for 69.1% of uploaders and produced videos with significantly higher quality than science communicators or institutional/patient accounts (all pairwise P < .05). Engagement indicators such as likes and shares were strongly correlated with each other (up to r = 0.92) but not with video quality (P > .05). TikTok and Bilibili play an increasing role in disseminating surgical knowledge about flap transplantation. Although TikTok videos were shorter, they demonstrated higher structural clarity and informational reliability. Videos uploaded by physicians showed the highest accuracy and transparency, while nonprofessional content exhibited lower educational value. These findings highlight the need for improved content regulation, professional participation, and evidence-based design to enhance the educational impact of medical short videos.
    Keywords:  Bilibili; TikTok; flap transplantation; health information quality; reconstructive surgery
    DOI:  https://doi.org/10.1097/MD.0000000000050873
  25. Lupus Sci Med. 2026 Sep 22. pii: e001912. [Epub ahead of print]13(2):
       OBJECTIVE: To analyse the content distribution and quality of lupus nephritis (LN)-related short videos on TikTok and Bilibili, compare differences across uploader types and examine the relationship between user engagement metrics and information reliability, thereby providing evidence to improve digital health communication strategies.
    METHODS: This cross-sectional study screened 250 LN-related videos from TikTok (n=143) and Bilibili (n=107). Basic video characteristics and uploader types were recorded. The Global Quality Score (GQS), modified DISCERN (mDISCERN; DISCERN, a validated instrument for evaluating the reliability of health information) and a content completeness instrument were used to assess video quality. All videos were independently evaluated by three senior rheumatology clinicians. Spearman correlation analysis was applied to evaluate associations between engagement indicators and quality scores.
    RESULTS: User engagement indicators (likes, comments, shares and favourites) and the proportion of professional uploaders were significantly higher on TikTok than on Bilibili (p<0.05). Although Bilibili videos were significantly longer, TikTok videos provided more comprehensive coverage of clinical manifestations and treatment (p<0.05). No significant difference in overall GQS was found between platforms (p>0.05), whereas TikTok achieved a slightly higher mDISCERN score (p<0.05). Professional institutions significantly outperformed non-professional uploaders across all quality indicators (p<0.05). Despite strong positive correlations among engagement indicators, major popularity metrics showed limited alignment with scientific quality scores.
    CONCLUSION: The quality of LN-related information on short-video platforms varies widely, with popularity metrics poorly reflecting scientifically assessed content quality. While TikTok shows a slight advantage in information reliability, both platforms require improvement in evidentiary support and content completeness. Professional institutions consistently produced higher-quality content, underscoring the need for broader participation by authoritative medical sources. A multistakeholder governance framework emphasising professional input, content quality standards and algorithm optimisation is essential to foster a trustworthy digital health ecosystem.
    Keywords:  Autoimmune Diseases; Lupus Nephritis; Prevalence
    DOI:  https://doi.org/10.1136/lupus-2025-001912
  26. Front Surg. 2026 ;13 1898401
       Background: Pancreaticoduodenectomy (PD), also known as the Whipple procedure, is a complex surgical procedure for pancreatic head and periampullary lesions. As short-video platforms increasingly shape the way patients and the public access health information, the quality and reliability of PD-related videos have important implications for patient education and public health communication.
    Objective: This study aimed to evaluate the source, content coverage, quality, reliability, and engagement characteristics of PD-related short videos on TikTok, Bilibili, and YouTube.
    Methods: Searches were conducted from January 26 to March 16, 2026, using newly registered accounts and Chinese and English keywords related to PD. After predefined inclusion and exclusion criteria were applied, 180 videos were included. Two independent raters with medical backgrounds assessed content coverage, the Global Quality Score (GQS), and the modified DISCERN (mDISCERN) score. Final GQS and mDISCERN scores were calculated as the average of the two raters' scores. Engagement indicators and uploader characteristics were extracted. Mann-Whitney U tests, Kruskal-Wallis H tests, chi-square tests, Spearman correlation analyses, interrater reliability analyses, and exploratory ROC and multivariable regression analyses were performed.
    Results: Interrater agreement was excellent for both GQS and mDISCERN. Among the 180 videos, 159 (88.33%) were uploaded by medical professionals and 21 (11.67%) by non-medical professionals. Medical-professional videos had significantly more likes and comments. In unadjusted comparisons, non-medical videos were longer and had higher GQS scores; however, uploader type was not independently associated with GQS after adjustment for duration and platform. mDISCERN scores did not differ significantly between the two groups. Most videos focused on management or procedural aspects of PD, whereas definitions, symptoms/indications, and outcomes/prognosis were less frequently addressed. Video duration was positively correlated with GQS and mDISCERN, and a sample-derived threshold of 160 s identified videos with GQS ≥4 with an AUC of 0.886. Significant differences in duration, engagement, quality scores, and uploader distribution were observed across platforms.
    Conclusion: PD-related videos on TikTok, Bilibili, and YouTube have some educational value but remain limited by incomplete content coverage and inconsistent reliability. Improving the completeness, transparency, comprehensibility, and platform governance of surgical health videos may help patients access more trustworthy information.
    Keywords:  Whipple procedure; health information; information quality; medical misinformation; pancreaticoduodenectomy; patient education; short video; social media
    DOI:  https://doi.org/10.3389/fsurg.2026.1898401
  27. Front Public Health. 2026 ;14 1980796
    Frontiers Production Office
      [This corrects the article DOI: 10.3389/fpubh.2026.1830047.].
    Keywords:  health communication; information quality; machine learning; short video; weight management
    DOI:  https://doi.org/10.3389/fpubh.2026.1980796
  28. Healthcare (Basel). 2026 Sep 09. pii: 2927. [Epub ahead of print]14(18):
      Background: People living with multiple sclerosis (MS) are increasingly turning to the Internet for information about their condition and its treatment. However, online health information varies widely in quality, and patients often lack the eHealth literacy needed to distinguish reliable from unreliable sources. Objective: This study examined how people with MS in the United Kingdom seek, evaluate, and judge the quality of online health information, with particular attention to information about medicines, and explored the features that they would value in a curated, quality-assessed information resource. Methods: A cross-sectional online survey was distributed via the MS Trust to adults with MS or their carers in the United Kingdom. The 55-item instrument, adapted from a previously validated questionnaire, captured demographic characteristics, Internet use, eHealth confidence, perceived quality indicators, assessment difficulties, and preferences for a curated information resource. Descriptive statistics summarised the sample, and Spearman's rank-order correlation tested associations between confidence, perceived importance of quality, and information-checking behaviours. Results: One hundred and fifty-two participants completed the survey. Because the number of individuals reached through the MS Trust's distribution channels was not recorded, a true response rate could not be calculated; the achieved sample represented approximately 38-40% of the a priori target of 400. Almost all participants (99%) used the Internet to find MS-related information, and 84% sought information about their medicines online. Recommendation by a healthcare professional was the strongest indicator of perceived information quality (17.8%), and MS specialists were the most trusted source overall (32.9%). However, 54.4% of participants expressed concerns about online information quality, and only 47.9% believed that search engines reliably returned high-quality websites. Confidence in evaluating online medicine information correlated strongly and positively with the perceived importance of information quality (ρ = 0.869, p < 0.001, n = 73) and with active checking behaviour (ρ = 0.677, p < 0.001, n = 73). Eighty-eight percent of participants endorsed the development of a single, quality-assessed website for MS-related medicine information, and 80% wished for visible details of how each source had been assessed. Conclusions: Despite high engagement with online information, people with MS remain uncertain about its quality and find independent assessment time-consuming and complex. A curated, transparently assessed information resource, ideally with clinician endorsement and visible quality indicators, was endorsed by the subset of participants who reached these items and represents a candidate direction for future digital health interventions in MS care. Because several preference items were answered by small subgroups (n = 35-79) drawn from an engaged charity membership, these findings are exploratory and require confirmation in more representative samples.
    Keywords:  United Kingdom; digital health; eHealth literacy; health information-seeking; information quality; multiple sclerosis; online health information; patient empowerment
    DOI:  https://doi.org/10.3390/healthcare14182927
  29. Front Public Health. 2026 ;14 1939476
       Background: Online health information-seeking behavior (OHISB) is important for nursing students, yet its heterogeneity and associations with digital health literacy and social support remain unclear.
    Objective: To identify OHISB profiles among nursing students and examine the hypothesized indirect association between digital health literacy and profile membership through perceived social support.
    Methods: This cross-sectional study included 772 nursing students. Data were collected using a sociodemographic questionnaire, the eHealth Literacy Scale, the Online Health Information-Seeking Behavior Scale, and the Perceived Social Support Scale. Latent profile analysis identified OHISB profiles. Multinomial logistic regression examined predictors of profile membership, with three-step analyses accounting for classification uncertainty. Structural equation modeling assessed indirect associations through perceived social support using bootstrap confidence intervals.
    Results: Among the nursing students, latent profile analysis identified three distinct online health information-seeking behavior profiles: high (n = 286, 37%), moderate (n = 328, 42.5%), and low (n = 158, 20.5%) OHISB profiles. Place of origin, family structure, household income, education level, grade, and social support were associated with profile membership (p < 0.05). In adjusted three-step analyses, higher perceived social support was associated with lower odds of high- (OR = 0.956) and moderate-OHISB (OR = 0.959) membership relative to low-OHISB membership. Digital health literacy was positively associated with perceived social support (B = 0.293, p < 0.001). Perceived social support was negatively associated with high- (B = -0.311, OR = 0.733, p < 0.001) and moderate-OHISB membership (B = -0.215, OR = 0.807, p < 0.001) relative to low-OHISB membership. Direct effects of digital health literacy were not significant (p = 0.637 and 0.511). Significant indirect effects were observed for high versus low (a × b = -0.091, 95% bootstrap CI: -0.136 to -0.060) and moderate versus low (a × b = -0.063, 95% bootstrap CI: -0.102 to -0.036) contrasts.
    Conclusion: Online health information-seeking behavior among nursing students was heterogeneous and characterized by three profiles. Perceived social support showed significant indirect associations between digital health literacy and OHISB profile membership. These findings highlight the potential importance of social support, while the cross-sectional design precludes causal inference.
    Keywords:  digital health literacy; latent profile analysis; nursing students; online health information-seeking behavior; social support; structural equation modeling
    DOI:  https://doi.org/10.3389/fpubh.2026.1939476