bims-skolko Biomed News
on Scholarly communication
Issue of 2026–07–26
37 papers selected by
Thomas Krichel, Open Library Society



  1. J Physiol. 2026 Jul 24.
      
    Keywords:  article publication charges; juniour researchers; open access; research funding
    DOI:  https://doi.org/10.1113/JP291774
  2. BMJ Glob Health. 2026 Jul 22. pii: e024444. [Epub ahead of print]11(7):
      
    Keywords:  Global Health
    DOI:  https://doi.org/10.1136/bmjgh-2026-024444
  3. Galen Med J. 2026 ;15 e3885
       Background: Article retraction means removing a published article from the journal because of ethical issues or scientific errors in order to correct the literature. In this study, we aimed to determine the reasons for retracting biomedical articles written by authors from Iran, Saudi Arabia, Pakistan, Egypt, and Turkey.
    Materials and Methods: This cross-sectional study included all retracted biomedical articles with first authors affiliated with Iran, Saudi Arabia, Pakistan, Egypt, or Turkey, retracted between September 1, 2010, and September 1, 2019. Data were extracted from Retraction Watch, MEDLINE, PubMed Central (PMC), Clarivate Analytics, and Scopus. Each article's information was entered into a data collection form and analyzed using SPSS version 24.
    Results: Of 436 retracted articles, Iran had the highest number (223), followed by Turkey (80), Egypt (72), Saudi Arabia (35), and Pakistan (26). Common causes of retraction included plagiarism, duplication, authorship issues, and fake peer review. In Iran, fake peer review (42.6%) and authorship issues (41.3%) were most prevalent. Significant inter- country differences were found in retraction frequency and causes. The most affected fields were biology, biochemistry, oncology, cardiovascular, surgery, and pathology.
    Conclusion: The results showed that scientific misconducts (plagiarism, duplication, authorship issues, and fake peer review) were the main reasons for retracting the articles in the five studied countries. To reduce such misconducts, regional regulatory policies, improved editorial practices, and enhanced research ethics training are urgently needed.
    Keywords:  Article Retraction; Authorship; Plagiarism; Retracted Article; Retractions; Scientific Misconduct
    DOI:  https://doi.org/10.31661/gmj.v15i.3885
  4. Am J Surg. 2026 Jul 17. pii: S0002-9610(26)00329-6. [Epub ahead of print]260 117144
      
    DOI:  https://doi.org/10.1016/j.amjsurg.2026.117144
  5. Front Artif Intell. 2026 ;9 1862980
       Introduction: Generative artificial intelligence (AI), particularly large language models, is rapidly becoming embedded in academic research and scholarly publishing. These systems assist with drafting, literature synthesis, and analytical writing, increasingly contributing to the production of academic text. This shift raises a central question: how should higher education institutions evaluate scholarly contribution when parts of research production become technologically mediated?
    Methods: This paper examines governance misalignment between publication standards and institutional evaluation systems in higher education. Drawing on an exploratory qualitative survey of 18 journal editors and associate editors across business-related disciplines, we analyze editorial perspectives on AI-assisted manuscript preparation, authorship, accountability, productivity, and academic evaluation.
    Results: We identify three recurring concerns: policy fragmentation, ambiguity surrounding authorship and accountability, and apprehension about AI-enabled productivity acceleration. We introduce the concept of metric distortion to describe the weakening relationship between measurable scholarly outputs and the intellectual labor those outputs are assumed to represent under conditions of AI-assisted production. As drafting and textual production become more technologically scalable, publication-based indicators may become less reliable proxies for conceptual contribution and interpretive effort.
    Discussion: Building on scholarship on academic capitalism, audit culture, and digital governance, the paper argues that generative AI functions as a stress test for existing evaluation regimes. We propose a process-based framework emphasizing disclosure, documented intellectual contribution, and institutional alignment between editorial governance and tenure evaluation systems.
    Keywords:  academic evaluation; authorship; distributed cognition; generative artificial intelligence; higher education; research integrity; scholarly governance
    DOI:  https://doi.org/10.3389/frai.2026.1862980
  6. J Empir Res Hum Res Ethics. 2026 Jul 24. 15562646261463812
      This study explores systemic factors that can influence unethical practices in academic publishing, particularly institutional pressures, editorial practices, and the impact of artificial intelligence (AI) tools. The study shifts responsibility away from individual actors. It takes a qualitative approach to examine how the interactions between researchers, editors, and institutional structures contribute to conditions in which unethical practices may occur. A qualitative research design was used to gather data through semi-structured interviews and observations with 30 participants, comprising university lecturers, postgraduate students, journal editors, and research support professionals from various regions.The results suggest that competition for publication and career advancement, as well as a lack of mentorship, create environments that participants see as encouraging unethical practices - such as author manipulation, use of third-party writers, and inappropriate use of AI - to occur. AI tools were specifically characterized as facilitating interventions that may support these practices in the context of systemic pressures. Furthermore, some participants expressed concerns about the editorial and financial arrangements, which could negatively affect the peer-review process.This study indicates that unethical publishing practices should be understood as resulting from complex systemic and cultural factors, rather than individual actions. The study identifies the need for institutional changes, guidance on AI use, and policies to promote research integrity. The qualitative nature and limited sample size of the study mean the results are explorative and suggestive of perceived patterns, rather than generalizable. We suggest that future research use mixed methods to investigate the extent and effects of these dynamics.
    Keywords:  Academic integrity; artificial intelligence; corruption in publishing; publication pressures; research ethics; unethical practices
    DOI:  https://doi.org/10.1177/15562646261463812
  7. J Clin Epidemiol. 2026 Jul 24. pii: S0895-4356(26)00305-7. [Epub ahead of print] 112429
       OBJECTIVE: To investigate reported paper mill involvement among retracted systematic reviews and meta-analyses (SR/MAs), and to compare retracted SR/MAs with and without reported paper mill involvement with respect to temporal patterns, disciplinary profile, country recorded in the database, and retraction characteristics.
    STUDY DESIGN AND SETTING: Cross-sectional study based on the Retraction Watch database. Records were screened to identify retracted SR/MAs. Reported paper mill involvement was assigned only when explicitly stated in the Retraction Watch record or associated retraction information; fake peer review was coded separately. Group comparisons used the Pearson chi-square and Mann-Whitney U tests. Associations with reported paper mill involvement were estimated using univariable Poisson regression with robust variance and reported as prevalence ratios (PRs) with 95% confidence intervals (CIs). These analyses were treated as exploratory, non-inferential internal contrasts within the dataset of already retracted SR/MAs.
    RESULTS: Among 69,163 retracted publications in the Retraction Watch dataset, 1,071 (1.5%) were identified as retracted SR/MAs. Of those, 191 (17.8%) were paper mill-associated and 880 (82.2%) were not. In the database country field, China was recorded for 732/1,071 retracted SR/MAs (68.3%) and 177/191 SR/MAs with reported paper mill involvement (92.7%). The annual number of paper mill-associated SR/MAs peaked in 2023 (n=152). Within this restricted dataset, paper mill-associated SR/MAs had a longer median time to retraction (511 vs 459 days; p=0.024), more frequent results-related concerns (89.0% vs 30.0%), data issues (53.4% vs 17.5%), referencing concerns (53.9% vs 13.6%), and ethical concerns (8.9% vs 2.6%) (all p<0.001), and less frequent multi-country authorship (8.9% vs 17.6%; p=0.004). In univariable Poisson regression analyses restricted to retracted SR/MAs, larger observed proportions of reported paper mill involvement were seen among records linked to China in the database country field (PR 5.86, 95% CI 3.45-9.93), more recent publication years (PR 1.20 per 1-year increase, 95% CI 1.16-1.23), the medical broad domain (PR 3.66, 95% CI 1.66-8.08), and records with results-related concerns (PR 11.88, 95% CI 7.68-18.39). These PRs should not be interpreted as likelihoods, risks, or rates beyond the analyzed Retraction Watch subset.
    CONCLUSION: In this Retraction Watch-based analysis, reported paper mill involvement was identified in 191/1,071 retracted SR/MAs (17.8%). These records were most often linked to China in the database country field (177/191; 92.7%) and clustered in recent years, particularly 2023. These findings should be interpreted as descriptive patterns within detected and retracted records, not as country-level rates, comparative performance indicators, or estimates of the true prevalence of paper mill activity.
    Keywords:  Evidence Synthesis; Peer Review; Research; Research Integrity; Retracted Publication; Scientific Misconduct
    DOI:  https://doi.org/10.1016/j.jclinepi.2026.112429
  8. Clin Child Psychol Psychiatry. 2026 Jul 23. 13591045261471231
      The rapid emergence of generative artificial intelligence (AI) has prompted growing debate about its role in academic writing and the implications for scholarly integrity. This editorial arose from an increasing number of manuscript submissions to the journal that appeared unusually polished, leading editors and reviewers to question whether AI-assisted writing had been used and, if so, whether this should influence editorial decisions. While AI has the potential to enhance clarity, accessibility, and efficiency, concerns remain about transparency, authorship, accountability, and the authenticity of scholarly contribution. To explore these issues, we conducted an experiment in which one section of this editorial was written entirely by a human co-author, while the second was developed using outputs from multiple generative AI tools that were subsequently substantially edited by the other human co-author, who retained responsibility for the final text. By juxtaposing these pieces, we invite readers to consider whether AI-assisted scholarship is distinguishable from human-authored writing, whether such distinctions matter, and how AI use might be disclosed. We argue that transparent reporting of AI assistance, rather than its prohibition, is essential to maintaining trust, accountability, and integrity in academic publishing.
    Keywords:  academic integrity; artificial intelligence; scholarship
    DOI:  https://doi.org/10.1177/13591045261471231
  9. World J Otorhinolaryngol Head Neck Surg. 2026 Jun 26.
       Objective: To compare the quality of scientific review articles generated by two artificial intelligence systems, ChatGPT and Gemini, with those written by human authors in the field of otolaryngology.
    Methods: Two otolaryngology topics, chronic rhinosinusitis and infantile subglottic hemangioma, were selected. For each topic, four AI-generated reviews (GPT-4.0 and Gemini 2.0; narrative and PRISMA-style) and one human-authored peer-reviewed review were included, yielding a total of 10 manuscripts (8 AI-generated, 2 human-authored). A blinded panel of seven board-certified otolaryngologists evaluated all manuscripts using a 5-point Likert scale across seven domains: scientific accuracy, depth of content, citation quality, structure and organization, readability and tone, critical insight, and overall scientific quality. Group comparisons were performed using linear mixed-effects models with random intercepts for reviewer and manuscript. Interrater reliability was assessed using Shrout-Fleiss intraclass correlation coefficients (ICC). Manual verification of AI-generated references was conducted to assess citation accuracy and fabrication.
    Results: Human-authored manuscripts received the highest ratings across all domains (overall quality 4.50 ± 0.76). GPT-4.0 demonstrated moderate performance (2.71 ± 1.46 overall), while Gemini 2.0 scored lowest (2.14 ± 1.01). Mixed-effects modeling demonstrated significant group differences across all domains (p ≤ 0.008). Citation quality showed one of the largest between-group differences and strong reliability [ICC(2,1) = 0.68; ICC(2,7) = 0.94]. Manual verification of 123 AI-generated references revealed high citation accuracy for GPT-4.0 (89.6% fully accurate; 0% fabricated) compared with Gemini 2.0 (71.7% fully accurate; 23.9% fabricated). Reviewers misclassified 50% of GPT-4.0 manuscripts as human-authored, correctly identified 93% of human-authored manuscripts, and classified 86% of Gemini manuscripts as AI-generated.
    Conclusion: GPT generated fluent, stylistically strong reviews but remained significantly inferior to human-authored manuscripts in analytical depth and citation integrity. Gemini 2.0 underperformed across all domains and demonstrated a substantial rate of fabricated citations. As large language models become integrated into academic workflows, transparent disclosure, structured fact-checking, and human oversight remain essential to safeguard scientific reliability.
    Keywords:  GPT‐4.0; citation integrity; generative artificial intelligence; otolaryngology; peer review; scientific writing
    DOI:  https://doi.org/10.1002/wjo2.70131
  10. World J Otorhinolaryngol Head Neck Surg. 2026 Mar 23.
       Objective: This study evaluated whether otolaryngologists can distinguish between human- and machine-written abstracts. The primary question was whether large language models (LLMs) produce abstracts comparable in clarity and usefulness to human-authored work, and whether reviewers can identify authorship with accuracy.
    Methods: A blinded cross-sectional design was used. Forty-eight abstracts were evaluated, consisting of twenty-four human-authored abstracts and 24 generated by four LLMs. Human abstracts were drawn from articles published after July 2025 to minimize overlap with LLM training data. Twenty otolaryngologists independently reviewed all abstracts. Using a structured rubric, raters classified authorship, rated clarity, usefulness, and confidence on 5-point scales, and provided optional free-text explanations. Group comparisons were performed using chi-square and Mann-Whitney tests, with Kruskal-Wallis tests for model-level analyses.
    Results: Overall recognition accuracy was 44.7%. Human-written abstracts were more often misclassified as AI than AI-generated abstracts were mistaken for human. Human abstracts received significantly higher clarity and usefulness scores than LLM abstracts, though effect sizes were small. Confidence did not correlate with correctness, indicating miscalibration of rater judgments. Model-level performance varied. Grok-generated abstracts were most easily identified as AI, whereas GPT-5 and Claude 3.5 more frequently resembled human writing. Free-text rationales commonly referenced style, vagueness, or lack of detail when AI authorship was suspected.
    Conclusion: LLMs generate abstracts that increasingly resemble human scientific writing, yet still lag in perceived usefulness and credibility. Clinicians were only moderately successful at detecting authorship and were frequently confident in incorrect classifications. These findings highlight both the promise and risks of AI-assisted scientific communication.
    Keywords:  abstract quality; authorship detection; large language models; otolaryngology; peer review; scientific publishing
    DOI:  https://doi.org/10.1002/wjo2.70104
  11. J Adv Med Educ Prof. 2026 Jul;14(3): 281-290
       Introduction: Modern large language models (LLMs) like ChatGPT (based on the GPT-4 architecture) and DeepSeek offer unprecedented capabilities for generating scientific text. However, their performance in replicating structured, high-quality scientific writing, especially compared to human-authored abstracts, remains insufficiently evaluated. To compare the abstract quality produced by human authors, ChatGPT/GPT4, and DeepSeek across: six evaluation criteria Clarity, Coherence, Conciseness, Accuracy, IMRaD Structure, and Language Quality, using blinded expert ratings and non-parametric statistical methods, specifically the Kruskal-Wallis test followed by pairwise Wilcoxon rank-sum tests with false discovery rate correction.
    Methods: We selected 23 medical and healthrelated research topics, each yielding three abstracts (human, ChatGPT, DeepSeek), for a total of 69 abstracts. Three raters scored each abstract. Kruskal-Wallis tests assessed group differences; Cliff's Delta (δ) was calculated as a nonparametric effect size for each comparison, suitable for ordinal data.
    Results: Across criteria, ChatGPT and DeepSeek significantly outperformed human authors in Clarity, Coherence, IMRaD Structure, and Language Quality. In contrast, Conciseness and Accuracy showed negligible effect sizes (|δ| <0.10), suggesting parity across all three sources.
    Conclusions: ChatGPT and DeepSeek achieved significantly higher scores in clarity, coherence, structure, and language quality, while showing comparable performance in conciseness and accuracy. These findings complement recent evaluations showing competitive medical and reasoning performance of DeepSeek models compared to proprietary LLMs. While shortform abstracts, expert oversight, and domain expertise remain critical, the results suggest that LLMs-particularly GPT4 and DeepSeek-can serve as effective tools in drafting scientific abstracts.
    Keywords:   Abstracting and indexing; Artificial intelligence; Deep learning; Medical writing; Natural language processing
    DOI:  https://doi.org/10.30476/jamp.2026.109761.2366
  12. Orthop Traumatol Surg Res. 2026 Jul 24. pii: S1877-0568(26)00222-7. [Epub ahead of print] 104801
       BACKGROUND: Artificial intelligence (AI) is increasingly integrated into scientific publishing workflows, yet no study has formally evaluated the ability of large language models (LLMs) to reproduce human editorial desk-review (R0) decisions in a general orthopaedic surgery journal. We investigated whether three commercially available LLMs could accurately replicate the editorial decisions of the Editorial Board of Orthopaedics & Traumatology: Surgery & Research (OTSR). The study addressed four questions: (1) Is the concordance between LLM and human R0 decisions satisfactory for editorial use? (2) Do LLMs exhibit a severity bias? (3) Do LLMs generate decision letters of acceptable quality, and do they reproduce the specific critiques of human reviewers? (4) Does prompt complexity influence LLM decision-making?
    HYPOTHESIS: LLMs used without task-specific fine-tuning or prior exposure to the study corpus would demonstrate at least moderate concordance (κ ≥ 0.40) with human editorial decisions.
    MATERIAL AND METHODS: A corpus of 32 manuscripts randomly selected from submissions to OTSR between 2025 and 2026 (n = 10 outright rejected at R0: 3 out of scope, 3 plagiarism/dual submission, 4 direct desk rejection; n = 11 accepted for peer review; n = 11 rejected after full peer review) was anonymised and independently evaluated, without task-specific fine-tuning or prior exposure to the study corpus, by ChatGPT (GPT-5.5, OpenAI), Gemini (3.1, Google), and Claude (Sonnet 4.6, Anthropic) using a structured prompt incorporating the OTSR guidelines. The primary outcome was assessed using Cohen's kappa between LLM and human binary decisions. Secondary outcomes included accuracy, inter-LLM agreement, domain-specific scoring, ARCADIA quality scoring of 115 eligible decision letters by two blinded raters with ICC, human-performed content concordance analysis, sensitivity analysis (structured vs. minimal prompt), and test-retest reproducibility at 24 hours.
    RESULTS: Overall accuracy (i.e. the decision was similar for LLM and editorial decision) was 59.4% for ChatGPT (19/32) and Claude (19/32), and 62.5% for Gemini (20/32). Cohen's kappa was near-zero for ChatGPT (κ = -0.05) and Gemini (κ = 0.00), and low for Claude (κ = 0.15). All LLMs showed systematic over-rejection of accepted manuscripts. No LLM identified plagiarism or simultaneous dual submission as a rejection motive. Test-retest concordance was 84.4 - 90.6% across models. ARCADIA quality scoring (n = 115 scorable letters, inter-rater ICC = 0.86, 95% CI 0.80-0.90) showed Claude achieved the highest scores (4.46 ± 0.32 /5), significantly above the human OTSR letters (4.01 ± 0.44, p < 0.001), ChatGPT (3.81 ± 0.49, p = 0.002), and Gemini (3.28 ± 0.49, p < 0.001). LLMs reproduced 30 - 41% of human reviewer-specific critiques, with Claude achieving the highest match (40.5%) without hallucinations. Switching to a minimal prompt markedly increased acceptance rates for Gemini (87.5%) and Claude (75%), while ChatGPT remained largely insensitive to prompt simplification (15.6%).
    CONCLUSION: The principal finding of this study is the dissociation between formal review quality and true editorial reliability. Although modern LLMs generated persuasive and methodologically structured decision letters, they failed to achieve meaningful concordance with real editorial decisions and displayed stable architecture-specific biases that were highly sensitive to prompt design. These results indicate that current LLMs reproduce the surface features of peer review more successfully than its underlying scientific and contextual reasoning. Consequently, LLMs may represent valuable supervised assistants for editorial workflows, but not reliable autonomous substitutes for human editorial expertise in orthopaedic scientific publishing.
    LEVEL OF EVIDENCE: IV; Observational pilot study, concordance analysis.
    Keywords:  AI editorial workflow limits (plagiarism-dual submission detection/contextual reasoning/human supervision).; Content concordance (human critiques reproduced/additional valid critiques/hallucinations); Editorial concordance (Cohen’s kappa/accuracy/sensitivity-specificity/over-rejection bias); LLM desk-review decisions (ChatGPT/Gemini/Claude/human editorial R0); Prompt sensitivity (structured vs minimal prompt/test-retest reproducibility); Review-letter quality (ARCADIA score/formal quality/synthetic peer review)
    DOI:  https://doi.org/10.1016/j.otsr.2026.104801
  13. J Neurosci Nurs. 2026 Jul 20.
       BACKGROUND: There has been an increase in the availability and use of artificial intelligence (AI) across many professional domains, and there are traces of AI in nearly every manuscript published using modern technology. It is becoming increasingly difficult for peer reviewers, editorial teams, and journal readers to identify the degree to which authors have used AI in the development of their manuscripts.
    METHODS: Our research group developed an AI scoring rubric to provide authors with an opportunity to self-disclose their use of AI.
    RESULTS: The Transparency and Reporting of Artificial Intelligence Contribution for Evaluating Submissions (TRACES) instrument provides a score from 0 to 40 across 3 domains: mechanics, writing, and illustrations. Higher scores indicate increased use of AI by the authors when writing or preparing a manuscript for submission.
    CONCLUSION: Authors in any field should self-report a TRACES score when submitting their manuscript. Journals may benefit from requiring authors to include a TRACES score when submitting a manuscript for peer review. While higher TRACES scores indicate greater use of AI, there is no specific cutoff provided to determine manuscript acceptance or rejection.
    DOI:  https://doi.org/10.1097/JNN.0000000000000903
  14. Br Dent J. 2026 Jul;241(2): 105-109
      One persistent issue undermining the integrity of scientific communication is spin - the misleading presentation or interpretation of results to make findings appear more favourable, significant, or impactful than they truly are, thereby leading to bias. This article examines the nature of spin in scientific communication and highlights the pressing need for increased awareness and preventive measures. Spin represents a form of reporting bias that can mislead readers, clinicians, and policymakers. These seemingly small distortions add up and contribute to biased scientific records. In a context where research underpins medical treatments, spin is not merely an academic or inconsequential practice - it can have significant real-world implications. Addressing spin bias is essential for preserving the integrity, credibility, transparency, and utility of research. Spin can be addressed through institutional reforms, enhanced editorial and peer review oversight, and training researchers, reviewers, and editors in transparent and responsible reporting. While the complete elimination of spin may be unrealistic, systematic and sustained efforts to identify, mitigate, and discourage it are essential for preserving the integrity of scientific literature. Fostering a research culture grounded in integrity, rigor, and accountability will ensure that scientific outputs genuinely contribute to knowledge advancement and serve the broader benefit of society.
    DOI:  https://doi.org/10.1038/s41415-026-9786-4
  15. Sci Rep. 2026 Jul 24. pii: 23202. [Epub ahead of print]16(1):
      Scientific research relies on transparent dissemination of data and its associated interpretations, including raw data, metadata, experimental design, and data processing details. Production and handling of research data represents an ongoing challenge, extending beyond publication into individual facilities, institutes and research groups, often termed Research Data Management (RDM). It is foundational to scientific discovery and aligned with the FAIR principles. Although the majority of peer-reviewed journals require raw data deposition in public repositories in alignment with FAIR principles, metadata frequently lacks standardization, hindering effective utilization and sharing of research findings. Here we present FRED, a generalized toolkit for FAIR metadata management in omics research based on a flexible, machine-readable YAML format. FRED enables (i) guided, dialog-based creation of metadata files, (ii) structured semantic validation, (iii) logical cross-file search, (iv) API-based integration with external systems, and (v) self-hosted web deployment. We demonstrate the utility of FRED through a complete annotation workflow applied to a published single-nucleus RNA-seq dataset, covering metadata generation, validation, repository-based discovery, and export to NCBI GEO submission format. FRED is designed for non-computational scientists and specialized facilities alike, and integrates into existing RDM infrastructure without requiring dedicated IT resources.
    DOI:  https://doi.org/10.1038/s41598-026-61886-9
  16. Nature. 2026 Jul 22.
      
    Keywords:  Careers; Machine learning; Publishing; Scientific community
    DOI:  https://doi.org/10.1038/d41586-026-02072-9
  17. Nature. 2026 Jul;655(8124): 1093-1094
      
    Keywords:  Machine learning; Publishing; Scientific community; Technology
    DOI:  https://doi.org/10.1038/d41586-026-02233-w
  18. Ann Intern Med. 2026 Jul 21.
      Prevalence and incidence are fundamental metrics with numerous applications in epidemiology. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guideline lacks specific items for reporting studies of disease prevalence or incidence. To address this gap, the STROBE Enhanced Prevalence and Incidence Criteria (STROBE EPIC) extension was developed in accordance with established methods for reporting guideline development. The authors generated an initial list of reporting items, conducted a modified Delphi process, and convened a face-to-face consensus meeting to confirm the need for a STROBE extension and to generate an early version of the checklist. They conducted 2 further Delphi surveys, first extending from typhoid and other invasive salmonelloses to all infectious diseases, and then to noncommunicable diseases and injuries. Finally, experts piloted the checklist on relevant manuscripts to critically assess if it was clear, concise, complete, and free of errors. An executive group curated the checklist after every survey round. The STROBE EPIC checklist comprises 47 items in the domains of title (1 item), abstract (2 items), introduction (1 item), methods (25 items), results (6 items), discussion (7 items), and other information (5 items). STROBE EPIC items address reporting of study design, adjustment factors for underreporting and underdiagnosis, denominator population estimation, case ascertainment methods, factors producing artefactual changes to observed disease prevalence or incidence, limitations of incomplete surveillance coverage, generalizability of short-duration studies, and data availability. The authors anticipate that the STROBE EPIC extension will be used by researchers, authors, modelers, burden-of-disease researchers, peer reviewers, and journal editors to optimize the presentation of epidemiologic evidence to support diverse health policy decisions.
    DOI:  https://doi.org/10.7326/ANNALS-25-02412
  19. J Int Soc Prev Community Dent. 2026 May-Jun;16(3):16(3): 272-278
       Objective: Systematic reviews (SRs) and meta-analyses (MAs) provide the highest level of scientific evidence for clinical recommendations and policy drafting in preventive dentistry. As readers frequently rely on abstracts for rapid appraisal, complete and transparent reporting is critical. The aim was to assess the reporting quality of abstracts of SRs and MAs published in international preventive dentistry journals using the Preferred Reporting Items for SRs and Meta-Analyses (PRISMA) 2020 extension for abstracts and to identify predictors of better reporting.
    Methods: This quality-of-reporting study analyzed 157 structured abstracts from seven purposively selected journals (quartiles 1-3) published between 2020 and 2022. Two independent, calibrated reviewers scored each abstract against the 12-item PRISMA 2020 checklist (0 = not reported, 1 = partially reported, 2 = fully reported). Inter-rater reliability was assessed using Cohen's kappa. Descriptive statistics and multiple linear regression were used for analysis.
    Results: Only the identification of the report as an SR/MA achieved 100% compliance across all journals. Moderate compliance was found for objectives and eligibility criteria. However, reporting was markedly suboptimal for synthesis methods, risk of bias, funding, and registration. Regression analysis revealed that a higher journal impact factor, larger journal volume, and affiliation with a professional society were significantly associated with higher reporting scores (P < 0.01).
    Conclusion: Adherence to PRISMA 2020 abstract guidelines in preventive dentistry remains suboptimal. Enhanced editorial oversight and strict adherence to reporting checklists are essential to improve the transparency and reproducibility of evidence used to shape clinical and policy decisions.
    Keywords:  Evidence-based dentistry; PRISMA 2020; meta-analysis; preventive dentistry; quality of reporting; systematic review
    DOI:  https://doi.org/10.4103/jispcd.jispcd_203_25
  20. Front Neuroimaging. 2026 ;5 1744569
      Increased racial and ethnic diversity in population neuroscience research is widely understood to facilitate better identification of subgroup effects and more generalizable findings. Consistency in reporting race and ethnicity population descriptor variables would allow the research community to better assess progress toward more representative datasets. One important lever for ensuring robust and consistent reporting of population descriptors are journal guidelines, and this review of current guidelines finds that there are opportunities for neuroscience journals to strengthen scientific rigor by more clearly delineating expectations with respect to reporting and operationalizing race and ethnicity population descriptors.
    Keywords:  ethics; ethnicity; journal publishing; neuroimaging; population descriptor; race
    DOI:  https://doi.org/10.3389/fnimg.2026.1744569
  21. J Ayurveda Integr Med. 2026 Jul 24. pii: S0975-9476(26)00092-6. [Epub ahead of print] 101408
      
    DOI:  https://doi.org/10.1016/j.jaim.2026.101408
  22. J Acquir Immune Defic Syndr. 2026 Jul 24.
      AI is a massively important tool for advancing the health of persons impacted by HIV. JAIDS is prioritizing rapid publication of research in this arena. AI also presents challenges to the quality of scientific publication and the bedrock peer review process that underpins scientific quality. JAIDS is committed to promoting the highest standard and rigor or science publication in the HIV space.
    Keywords:  AI and Research; Artificial Intelligence; HIV; Peer Review; Publishing
    DOI:  https://doi.org/10.1097/QAI.0000000000003943
  23. Arthroscopy. 2026 Jul 23.
      Arthroscopy aims to be the primary home for musculoskeletal biologics research. To this end, we introduce our Fourth Annual Musculoskeletal Biologics Special Issue with Guest Editor Brian Cole, M.D., whose team selected the top publications from the past year in the Arthroscopy Family of Journals. We reiterate our Call for Papers of clinically relevant and high-impact musculoskeletal biologics research. We further highlight our lead article and Expert Opinion from top authors outlining high-yield topics identified by the 2026 Biologic Research Think Tank with eager anticipation to publish future literature of nascent biologic substances harnessed to augment musculoskeletal treatment.
    DOI:  https://doi.org/10.1002/arj.70487
  24. Acad Med. 2026 Jul 21. pii: wvag225. [Epub ahead of print]
      
    Keywords:  academic medicine; medical education scholarship; scholarly publishing
    DOI:  https://doi.org/10.1093/acamed/wvag225
  25. Acad Med. 2026 Jul 24. pii: wvag230. [Epub ahead of print]
      
    Keywords:  academia; diversity; equity; inclusion; organizational innovation; research
    DOI:  https://doi.org/10.1093/acamed/wvag230
  26. J Korean Med Sci. 2026 Jul 20. 41(28): e286
      Misuse of statistical methods in biomedical research remains widespread, undermining scientific integrity and public health. Flawed analyses can lead to misleading clinical guidelines, unnecessary treatments, or concealment of true treatment effects. Statistical errors often result from a structural training gap that affects the research process from study design to reporting, rather than from individual incompetence. Authors should be able to select appropriate measures of central tendency based on the data distribution, apply parametric and nonparametric tests correctly, use regression analyses with attention to statistical power and collinearity, and interpret P values in the context of confidence intervals. The growing number of retractions in statistics and the continued citation of retracted articles highlight the limitations of current peer-review processes. Statistical Package for the Social Sciences is the most widely used statistical software in health sciences research, while R/RStudio is gaining increasing adoption due to its open-source nature, analytical flexibility, and capacity for fully reproducible and transparent analyses. Software selection can significantly influence analytical outcomes due to differences in default algorithms. Therefore, full disclosure of software versions and procedures is essential for reproducibility. Integrating artificial intelligence-based tools into statistical workflows introduces risks, including output hallucination, a lack of algorithmic transparency, and unresolved accountability. These tools should not be used without internationally accepted validation guidelines. At the publisher level, employing statistical editors, enforcing open data-sharing policies, and adopting contributor role taxonomies such as Contributor Roles Taxonomy are the most effective structural interventions. Enhancing the statistical quality of biomedical literature requires coordinated, sustained efforts from all stakeholders in scientific communication.
    Keywords:  Artificial Intelligence; Biostatistics; Data Interpretation, Statistical; Editorial Policies; Peer Review; Periodicals as Topic; Retraction of Publication
    DOI:  https://doi.org/10.3346/jkms.2026.41.e286
  27. Nature. 2026 Jul;655(8124): 823
      
    Keywords:  Authorship; Education; Machine learning; Publishing
    DOI:  https://doi.org/10.1038/d41586-026-02227-8