bims-aukdir Biomed News
on Automated knowledge discovery in diabetes research
Issue of 2026–09–20
nineteen papers selected by
Mott Given



  1. 7th Int Conf Wirel Intell Distrib Environ Commun (2024). 2025 ;237 109-122
      Diabetes, a severe and chronic condition characterized by elevated blood glucose levels, has been a significant health challenge. In recent years, Machine learning has shown promise in predicting the identification of diabetes and its related complications. However, the development of these predictive models has been hindered by inconsistencies, poor data quality, inappropriate correlational models, and the inherent complexity of clinical data. These issues often render existing machine learning algorithms for diagnosis and treatment ineffective. In this paper, we propose a novel survey-based machine-learning model for the early detection of diabetes. This model, which has the potential to enhance the effectiveness of existing machine learning algorithms significantly, utilizes the synthetic minority oversampling technique and machine learning techniques, thereby advancing healthcare technology and aiding healthcare professionals in diabetes diagnosis and creating effective prevention strategies.
    Keywords:  Clinical data analysis; Diabetes prediction; Healthcare; Machine learning; Preventive care
    DOI:  https://doi.org/10.1007/978-3-031-80817-3_7
  2. Front Endocrinol (Lausanne). 2026 ;17 1902508
       Background: Diabetic nephropathy (DN) and diabetic retinopathy (DR) are common microvascular complications of type 2 diabetes mellitus (T2DM) and may require different diagnostic and management pathways. This study aimed to develop and validate an interpretable machine learning model based on routine laboratory data to differentiate prevalent DN from prevalent DR among hospitalized patients with type 2 diabetes.
    Methods: Data were collected from a large tertiary hospital in China and split into a training/internal validation cohort (DN: 2,309 cases; DR: 855 cases) and an independent held-out validation cohort (DN: 578 cases; DR: 214 cases). A total of 47 routinely available laboratory and demographic variables were extracted from electronic health records (EHRs). Seven machine learning algorithms were developed and compared, with recursive feature elimination (RFE) employed to identify the most informative subset of features and enhance model performance and interpretability. Model discrimination was assessed using the area under the receiver operating characteristic curve (AUC) and the area under the precision-recall curve (AP), while SHAP values were used to interpret feature importance and explain individual-level predictions.
    Results: The extreme gradient boosting (XGBoost) classifier demonstrated the highest predictive performance among the seven machine learning algorithms evaluated. After selecting the top five features based on importance rankings, an explainable XGBoost model was constructed. This final model achieved strong apparent discrimination in both the training/internal validation cohort (AUC = 0.991, 95% CI: 0.989-0.994; AP = 0.979, 95% CI: 0.973-0.984) and the held-out validation cohort (AUC = 0.997, 95% CI: 0.996-0.999; AP = 0.993, 95% CI: 0.988-0.997). SHAP analysis further identified α-hydroxybutyrate dehydrogenase, creatine kinase-MB, creatinine, urinary α1-microglobulin, and N-acetyl-β-D-glucosaminidase as the most influential features contributing to complication risk prediction.
    Conclusions: An explainable machine learning model for predicting complications in patients with T2DM demonstrated high feasibility and effectiveness, indicating strong potential to support clinical management and improve patient outcomes. By incorporating SHAP analyses, the model addresses key concerns regarding transparency and clinical decision-making. These findings highlight the model's potential for real-world clinical implementation.
    Keywords:  SHAP; diabetic complications; machine learning; prediction model; type 2 diabetes mellitus
    DOI:  https://doi.org/10.3389/fendo.2026.1902508
  3. Diagnostics (Basel). 2026 Sep 07. pii: 2882. [Epub ahead of print]16(17):
      Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a unified experimental framework remain limited. This study systematically compares representative standard and lightweight CNN- and transformer-based architectures for multi-class DR grading. Methods: Six ImageNet-pretrained deep-learning models, including ResNet50, EfficientNet-B0, MobileNetV2, Vision Transformer (ViT), Swin-Tiny, and Swin Transformer, were evaluated on the APTOS 2019 retinal fundus image dataset under a unified experimental configuration with consistent preprocessing, data augmentation, training, and evaluation settings. All models were fine-tuned and evaluated independently over five runs with different random seeds. Their performance was assessed using accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve (AUC), Quadratic Weighted Kappa (QWK), per-class analysis, computational efficiency, and statistical analysis. Results: Transformer-based models generally achieved higher mean classification performance than the evaluated CNN-based models. Swin-Tiny achieved the highest mean accuracy (82.3%), macro F1-score (64.4%), weighted F1-score (82.1%), and QWK (89.8%) across the five runs. Among the CNN-based models, EfficientNet-B0 achieved the strongest overall classification performance, whereas MobileNetV2 provided the lowest computational complexity. The results also highlighted differences in learning behavior and computational requirements across the evaluated architectures. Repeated experiments demonstrated stable performance across different random seeds, supporting the reliability of the proposed evaluation. Conclusions: Overall, this study provides a comprehensive comparison of representative CNN- and transformer-based architectures under consistent experimental settings and offers practical guidance for selecting suitable deep learning models for automated diabetic retinopathy screening.
    Keywords:  convolutional neural networks; deep learning; diabetic retinopathy; fundus imaging; multi-class classification; transfer learning; transformer-based models
    DOI:  https://doi.org/10.3390/diagnostics16172882
  4. Curr Drug Discov Technol. 2026 Sep 08.
      Diabetic retinopathy (DR) is a condition that progressively affects the microvasculature, commonly seen in individuals with diabetes, which is a leading cause of visual impairment across the globe. Approximately one-third of people with diabetes will eventually develop DR, and it remains a public health issue because the retinal damage is irreversible if the disease continues to go undiagnosed and untreated. Although various screening methods exist, including fundus photography and optical coherence tomography (OCT), they are often limited as a result of availability and expertise required for diagnosis and management. Recent advancements in artificial intelligence (AI) are transforming DR screening and diagnosis. AI has demonstrated the potential to automate retinal image evaluation with accuracies that, in some studies, approach or exceed existing standards of care; however, reported performance varies across algorithms, datasets, and clinical settings, highlighting the need for cautious interpretation and further validation in diverse populations. Machine learning (ML) & deep learning (DL) models, particularly convolutional neural networks (CNNs), have been found to have excellent performance in identifying DR-related abnormalities, including mild, moderate, and severe DR, and diabetic macular edema, at sensitivities and specificities exceeding traditional screening. AI tools under various names, such as IDx-DR, EyeArt, and Google's AI algorithm, are quick and price-effective tests using an AI solution for early diagnosis of DR in either developed or developing economies. In addition, AI can also support predictive analytics, risk stratification, and treatment decision-making, so that ophthalmologists can provide patients with better management options. Importantly, AI also comes with its own challenges, such as data harmonization and regulatory approval for deploying AI in clinical care. This review will focus on the existing literature on the applications of AI in diagnosing and management of DR and will outline the importance of this innovation.
    Keywords:  Diabetic retinopathy; artificial intelligence; diabetes mellitus; ophthalmology
    DOI:  https://doi.org/10.2174/0115701638445311260825100649
  5. Int J Retina Vitreous. 2026 Aug 14. pii: 129. [Epub ahead of print]12(1):
       BACKGROUND: Diabetic Retinopathy (DR) is a leading cause of vision impairment and a major long-term microvascular complication of diabetes. Early detection and prevention of diabetes-related complications through advanced imaging and software-assisted patient management remain important clinical priorities. Microaneurysms (MAs) are the earliest and most subtle indicators of DR, but their small size and low contrast often lead to missed detection during manual fundus examination, delaying intervention. Automated MA segmentation is therefore essential for large-scale DR screening.
    METHODS: We conducted a controlled empirical evaluation of localized patch-based training for microaneurysm (MA) segmentation, benchmarking a compact, imbalance-aware U-Net model including learnable transposed-convolution up-sampling against a full-image U-Net trained on identical data. Fundus images and corresponding MA masks were divided into non-overlapping 256 × 256 patches, increasing MA pixel density per training sample by approximately 60-fold relative to full-image input thereby directly targeting the extreme class imbalance that causes full-image models to collapse to all-background predictions. The model was trained and evaluated on the publicly available IDRiD [Indian Diabetic Retinopathy Image Dataset] and DDR [Dataset for Diabetic Retinopathy] datasets. Performance was assessed using Intersection over Union (IoU), Dice coefficient, accuracy, recall, and precision.
    RESULTS: Performance was evaluated on a representative held-out subset of 100 images selected from an independent pool of 514 MA-annotated test images spanning the IDRiD and DDR datasets. The patch-based model achieved an overall pixel-level accuracy of 99.89%, a Dice coefficient of 76.9%, IoU of 63.2%, recall (sensitivity) of 69.1% and precision of 88.7%, while the full-image U-Net trained on the same data failed to recover any MA pixels (IoU = 0, Dice = 0) despite comparable pixel accuracy thereby demonstrating that overlap-based metrics, not accuracy, are the decisive criterion for this task. Per-image lesion-coverage analysis showed a mean match rate of approximately 60% of annotated MA contours, providing a clinically interpretable read-out beyond pixel overlap.
    CONCLUSION: This controlled empirical evaluation demonstrates that localized patch-based training is an effective and computationally efficient strategy for overcoming the extreme class imbalance that causes conventional full-image U-Nets to systematically miss microaneurysms. The proposed patch-based U-Net showed promising microaneurysm segmentation performance on a held-out subset of IDRiD and DDR images; further evaluation on the additional datasets and independent external cohorts can aid in the clinical screening utility.
    Keywords:  Diabetic retinopathy; Early DR; Fundus imaging; Medical image analysis; Microaneurysm segmentation; Patch-based deep learning; U-Net
    DOI:  https://doi.org/10.1186/s40942-026-00919-x
  6. PLoS One. 2026 ;21(9): e0352313
      Diabetes is a chronic disease that significantly increases the risk of serious complications such as cardiovascular disorders and kidney failure. Early detection through predictive modeling can lead to timely interventions and significantly improve patient health outcomes. Several machine learning approaches have been proposed for predicting diabetes, but the main focus has been on improving prediction accuracy, while interpretability has received limited attention. To address this gap, we present a robust and explainable machine learning framework based on a stacked ensemble model that uses Random Forest, Support Vector Machine, and Gradient Boosting as base learners and Catboost as the meta-learner. The model was trained on the PIMA Indians Diabetes dataset using a preprocessing pipeline that included standard scaling, analysis of variance (ANOVA)- F-score-based feature selection, and class balancing with the synthetic minority oversampling technique (SMOTE). The proposed ensemble model outperformed the latest methods with an accuracy of 86%. We integrated explainable AI techniques such as Local Interpretable Model-Agnostic Explanations (LIME) and Shapley Additive Explanation (SHAP) to enhance transparency, which provide both local and global interpretability by identifying the most influential features contributing to each prediction, thus supporting more informed and trustworthy decision-making in healthcare applications.
    DOI:  https://doi.org/10.1371/journal.pone.0352313
  7. Patient Prefer Adherence. 2026 ;20 624754
       Background: Physical activity adherence is critical for effective diabetes management, yet non-adherence remains common in community settings. This study sought to construct and internally evaluate an interpretable machine learning model for identifying poor physical activity adherence among patients with Type 2 diabetes mellitus (T2DM) in a community-based setting.
    Methods: This single-center cross-sectional study included 207 patients with T2DM receiving community-based care at the Shiqiao Community Health Service Center, China, between January and September 2025. Candidate predictors were selected using Elastic Net regularization followed by multivariable logistic regression. Multiple machine learning algorithms were developed and compared using a randomly allocated training dataset (70%, n = 145) and an internal validation dataset (30%, n = 62). Model discrimination and clinical utility were assessed using the area under the receiver operating characteristic curve (AUC) and decision curve analysis (DCA), respectively. Shapley Additive exPlanations (SHAP) was used to interpret individual predictor contributions to model outputs.
    Results: Among the 207 participants, 113 (54.6%) demonstrated high physical activity adherence. The final model incorporated five key predictors: Summary of Diabetes Self-Care Activities (SDSCA) score, estimated glomerular filtration rate (eGFR_CKD-EPI_2009), family history of diabetes, International Physical Activity Questionnaire (IPAQ) score, and albumin (ALB). The SVM model achieved the highest discrimination ability among the tested algorithms and was subsequently selected as the final model (AUC = 0.787, 95% CI: 0.671-0.903). SHAP analysis provided transparent visualization of the relative contribution of each predictor to model outputs.
    Conclusion: This study provides an explainable machine learning approach for identifying key variables associated with physical activity adherence among individuals with T2DM in community diabetes management. Further validation in independent populations is needed before the model can be considered for risk assessment and personalized behavioral intervention planning in community diabetes care.
    Keywords:  explainable artificial intelligence; machine learning; physical activity adherence; predictive modeling; risk stratification; type 2 diabetes mellitus
    DOI:  https://doi.org/10.2147/PPA.S624754
  8. Diagnostics (Basel). 2026 Sep 01. pii: 2813. [Epub ahead of print]16(17):
      Background/Objectives: Diabetes mellitus is a chronic metabolic disorder characterized by impaired regulation of blood glucose due to defects in insulin secretion, insulin action, or both. Physiological and lifestyle factors vary among individuals. General medicine is not applicable to all patients. In this scenario, personalized medicine for each individual becomes costly. Effective management of continuous glucose levels with accurate insulin dosage is challenging. To overcome this, a digital twin (DT)-based insulin dosage simulator with an individual's metabolic system is proposed in this work. Methods: Various machine learning techniques, mathematical models of physiology, and risk assessment using probability are used to predict the dynamics of patient-specific glucose-insulin. Parameters such as carbohydrate intake, sleep patterns, medications, and physical activity were incorporated into this model to capture real-world variations in daily life. For glucose-insulin interactions, the Bergman Minimal Model (BMM) is used; for time-of-day variability, a circadian insulin sensitivity model is used; and for predicting metabolic risks, Bayesian risk estimation (BRE) is used, which includes hyperglycemia risk. To enhance transparency and interpret model predictions, explainable artificial intelligence (XAI) methods are employed. Results: The simulation results showed improved glucose prediction accuracy, enhanced detection of hypoglycemia risk, and optimized insulin dosing strategies compared with traditional approaches. Conclusions: Overall, the proposed digital twin model offers a scalable solution using the latest techniques A "Prescriptive Analytical Framework" is provided using the BMM and BRE for personalized diabetes management and decision support for clinicians.
    Keywords:  Bayesian risk estimation (BRE); Bergman minimal model (BMM); digital twin (DT); explainable AI (XAI); hyperglycemia; hypoglycemia
    DOI:  https://doi.org/10.3390/diagnostics16172813
  9. Front Endocrinol (Lausanne). 2026 ;17 1905144
       Introduction: Although observational studies have established associations between three major liver enzymes (alanine aminotransferase (ALT), aspartate aminotransferase (AST), and gamma-glutamyl transferase (GGT)) and type 2 diabetes (T2D), their ability to improve T2D identification using machine learning (ML) approaches remains underexplored. This study aimed to develop and validate ML-based diagnostic models for identifying T2D by incorporating these liver enzymes.
    Methods: Data from two independent cohorts were analyzed: the US National Health and Nutrition Examination Survey (N = 15,528) and a Chinese health examination database (N = 4,952). Twelve demographic and biochemical features were combined with liver enzymes to train eight ML models, including logistic regression, support vector machine, random forest, K-nearest neighbors, classification and regression trees, gradient boosting decision tree, LightGBM, and XGBoost. Model performance was assessed using standard metrics including accuracy, precision, recall, F1-score, and area under the curve (AUC).
    Results: Incorporating liver enzymes consistently improved model performance across the algorithms. The XGBoost model showed particularly strong performance, with baseline metrics (accuracy = 0.696, precision = 0.314, F1 = 0.450, AUC = 0.803) increasing to 0.743, 0.346, 0.468, and 0.814, respectively, after inclusion of liver enzymes. Comparable improvements were observed across algorithms and in both the US and Chinese cohorts.
    Conclusions: ML models incorporating routinely measured liver enzymes improve T2D identification across US and Chinese datasets. These findings indicate the potential utility of liver enzymes as accessible adjunctive indicators for diabetes risk stratification and improved T2D identification.
    Keywords:  NHANES; liver enzymes; machine learning; prediction model; type 2 diabetes
    DOI:  https://doi.org/10.3389/fendo.2026.1905144
  10. Front Artif Intell. 2026 ;9 1795751
       Introduction: Diabetic foot ulcer (DFU) image classification can support timely clinical triage; however, many existing systems provide limited severity granularity, rely on single-source image datasets, and predominantly use pixel-level explanation methods. This study developed a seven-class, severity-aware, multi-task artificial intelligence framework for interpretable DFU image analysis and smartphone-class deployment.
    Methods: The framework combines domain-adaptive SimCLR pretraining, U-Net-based image refinement, EfficientNet-B0 feature extraction, a lightweight windowed transformer fusion block, and parallel classification and ordinal severity-regression heads. Four public data sources-DFUC 2020/2021, AZH, Medetec, and an IWGDF-aligned ordinal repository-were harmonized and deduplicated, yielding an analytical corpus of 24,925 original images. Evaluation included a 3,739-image held-out test set, five-fold stratified cross-validation, leave-one-source-out testing, an independent prospective clinician-audited cohort of 366 images, calibration and uncertainty analyses, concept-based interpretability assessment, and mobile-device profiling.
    Results: On the held-out test set, the framework achieved 95.5% accuracy (95% CI, 94.7-96.1) and a macro F1 score of 0.953, while mean five-fold cross-validation accuracy was 96.5% ± 0.3 percentage points. External ordinal severity estimation yielded a Spearman correlation of ρ = 0.91. In the clinician-audited cohort, weighted Cohen's κ for model-clinician severity agreement was 0.952. Model-assisted review was associated with a change in recorded management intent in 27.6% of cases, and 92.6% of generated explanations were rated as clinically useful. Full on-device inference, including gradient-weighted class activation mapping and concept-based explanation generation, required 322 ms per image on a Snapdragon 8 Gen 1 device.
    Discussion: The proposed framework demonstrates how multi-source representation learning, multi-task severity modeling, uncertainty assessment, concept-based interpretability, and efficient edge deployment can be integrated within a unified AI pipeline for DFU image analysis. The findings indicate high classification performance, strong ordinal severity agreement, clinically interpretable outputs, and technically feasible smartphone-class inference. Further multicenter prospective validation across diverse populations and acquisition conditions is required before translation into routine clinical decision-support settings.
    Keywords:  diabetic foot ulcer; explainable artificial intelligence; external validation; hybrid CNN–transformer; mobile deployment; multi-task learning; self-supervised learning; severity classification
    DOI:  https://doi.org/10.3389/frai.2026.1795751
  11. J Clin Med. 2026 Aug 24. pii: 6542. [Epub ahead of print]15(17):
      Background: Achieving remission of type 2 diabetes mellitus (T2DM) after bariatric surgery represents a critical opportunity to reduce long-term diabetes-related complications, including cardiovascular disease, nephropathy, neuropathy, and retinopathy. However, remission rates vary widely across patients, and identifying modifiable clinical and behavioral determinants remains essential for optimizing integrated metabolic care. Objectives: In the current study, we aimed to (1) classify type 2 diabetes mellitus (T2DM) remission status after bariatric surgery through clinical, anthropometric, and behavioral variables at follow-up; (2) identify the main model-based determinants of remission status and explainable machine learning using the preoperative model for baseline risk stratification with surgical candidates. Methods: We performed a retrospective cross-sectional study on 233 patients with T2DM who had bariatric surgery at a tertiary referral center. We made use of two analytical frameworks: a full-feature approach to identify the current remission status in a cross-sectional manner and a preoperative approach to make a temporal classification of the baseline for the first time. We trained and internally assessed 14 machine learning and deep learning classifiers. We evaluated model interpretability using SHAP. Results: In the full-feature cross-sectional classification, the Bottleneck Network performed best (ROC AUC = 0.889). SHAP data identified percentage weight regain, pre- and post-surgical body mass index, HbA1c, and oral hypoglycemic agent use as the dominant model-associated factors. In the restricted preoperative setting, the Extra Trees model achieved an AUC of 0.707, which is a lower level but still represents good baseline risk stratification performance. Conclusions: The findings indicate that remission status after bariatric surgery is both clinical as well as behavioral, but the post-operative or contemporaneously assessed variables should be looked at as classification (as opposed to prediction) models. The preoperative model may be able to be used in risk stratification, but clinical validation and prospective evaluation should be made prior to clinical implementation.
    Keywords:  artificial intelligence; bariatric surgery; diabetes remission; integrated diabetes care; machine learning; metabolic complications; risk stratification; type 2 diabetes mellitus
    DOI:  https://doi.org/10.3390/jcm15176542
  12. Front Med Technol. 2026 ;8 1908856
       Problem: While voice-based analysis has emerged as a potential non-invasive approach for exploring acoustic alterations associated with Type 2 Diabetes Mellitus (T2DM), current approaches often fail to fully capture disease-relevant vocal patterns due to simplified recording protocols and insufficient feature modeling.
    Aim: This study aims to investigate the feasibility of voice-based classification of T2DM using a multi-feature fusion framework, leveraging sustained vowel phonations and integrating multiple acoustic feature types to enhance diagnostic performance.
    Methods: Voice recordings were collected from 378 participants, including T2DM patients and healthy controls. Each participant pronounced six Mandarin vowels, from which acoustic features-including Mel-frequency cepstral coefficients (MFCCs), glottal parameters, and eGeMAPS descriptors-were extracted. Neural networks, including convolutional and deep architectures, were applied to capture pathology-relevant segments. Vowel-level fusion modules aggregated features across vowels, and an attention-based hierarchical fusion mechanism combined the three feature types, adaptively emphasizing features most relevant to T2DM.
    Results: The proposed dual-branch hierarchical fusion model achieved 80.48% detection accuracy.
    Conclusion: These results suggest that sustained vowel phonations contain potentially informative acoustic patterns associated with T2DM status. However, further validation using independent cohorts is required before clinical application.
    Keywords:  T2DM detection; convolutional neural networks; deep neural networks; non-invasive T2DM; speech features
    DOI:  https://doi.org/10.3389/fmedt.2026.1908856
  13. PLOS Digit Health. 2026 Sep;5(9): e0001734
      Artificial intelligence can already automate selected diabetes tasks, but the evidence is uneven. Automated insulin delivery is established in type 1 diabetes; in type 2 diabetes, validated screening and insulin-adjustment tools support a tiered, human-governed model rather than wholesale autonomy.
    DOI:  https://doi.org/10.1371/journal.pdig.0001734
  14. Ophthalmol Sci. 2026 Oct;6(10): 101335
       Objective: OCT angiography (OCTA) images present challenges for clinical and research use due to variability and noise. The aim was to develop artificial intelligence models for OCTA image quality evaluation using the Artificial Intelligence Ready and Equitable Atlas for Diabetes Insights (AI-READI) data set.
    Design: Cross-sectional study.
    Subjects: Six thousand two hundred sixty-nine OCTA 6 × 6-mm and 325 12 × 12-mm macula-centered photographs of the superficial vascular plexus from 1067 AI-READI study participants, and 1100 6 × 6-mm OCTA photographs from 539 patients from the Massachusetts Eye and Ear Infirmary (MEEI).
    Methods: All photographs were labeled as acceptable or poor quality by two ophthalmologists and five medical students. Predefined training and validation sets were used to fine-tune four deep learning (DL) models (ResNet, EfficientNet, Vision Transformer [ViT], and ConvNeXt v2) via hyperparameter grid search.
    Main Outcome Measures: Model performance on all photographs from the AI-READI and MEEI test sets were assessed by accuracy, sensitivity, specificity, and area under the receiver operating characteristic (AUROC) scores.
    Results: The accuracy/AUROC of the fine-tuned ResNet, EfficientNet, ViT, and ConvNeXt v2 models on the 6 × 6-mm AI-READI test set were 83.5%/0.904, 82.5%/0.904, 83.5%/0.908%, and 82.9%/0.913, with sensitivities/specificities of 77.9%/86.1%, 73.4%/89.9%, 69.2%/92.7%, and 79.8%/84.2%, respectively. The accuracy/AUROC of these models on the 6 × 6-mm MEEI test set were 91.5%/0.925, 94.3%/0.974, 93.9%/0.974%, and 88.0%/0.957, with sensitivities/specificities of 73.2%/98.4%, 84.5%/97.3%, 86.8%/97.2%, and 84.4%/92.7%, respectively. The accuracies/AUROC for the 12 × 12-mm photographs were 72.3%/0.775, 76.4%/0.775, 66.2%/0.979%, and 83.7%/0.834, with sensitivities/specificities of 95.7%/54.5%, 95.7%/57.8%, 97.9%/42.2%, and 96.4%/74.1%, respectively.
    Conclusions: The AI-READI data set contains heterogeneous data usable for high-performing DL models for OCTA image quality assessment. These models were generalizable across institutions, cameras, and image sizes and may be used to rapidly screen OCTA image quality for use in future clinical studies.
    Financial Disclosures: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
    Keywords:  Artificial intelligence; Deep learning; Image quality.; Imaging; OCT angiography
    DOI:  https://doi.org/10.1016/j.xops.2026.101335
  15. Diagnostics (Basel). 2026 Aug 25. pii: 2707. [Epub ahead of print]16(17):
      Background/Objectives: Double reading with arbitration improves diagnostic reliability in fundus screening but requires repeated human interpretation. Whether artificial intelligence (AI) can serve as a first-pass decision source within this workflow, and whether this applies consistently across retinal diseases with differing inter-reader agreement, remains unclear. Methods: In this retrospective study, 6904 color fundus photographs from 2593 patients at a tertiary screening center were analyzed for age-related macular degeneration (AMD), diabetic retinopathy (DR), and retinal vein occlusion (RVO). Three retina specialists independently labeled each image, and an AI system provided binary classifications at a prespecified operating threshold targeting 0.99-sensitivity. In AI-human double reading, the AI and one reader independently interpreted each image, and a second reader arbitrated discordant cases; this was compared with conventional human-human double reading. Three reader combinations were evaluated per disease. Results: AI-human reading required 1.02-1.11 human reads per image versus 2.00-2.10 for human-human reading. For DR and RVO, human-human reading yielded sensitivities of 0.974 and 0.982 and specificities of 0.999 and 1.000, respectively. Across AI-human combinations, sensitivity and specificity did not differ significantly from human-human reading (all p ≥ 0.05; specificity differences ≤0.001). For AMD (human-human sensitivity 0.845, specificity 1.000), AI-human sensitivity varied: two combinations were higher (0.916 and 0.950; both p < 0.001) and one comparable (0.842; p = 0.742). AMD specificity remained ≥0.978. Conclusions: AI-human reading halved human reading volume without significant loss for DR and RVO; AMD varied by configuration. AI use within double-reading workflows should account for disease-specific inter-reader agreement, reader composition, and operating threshold. This was a single-center, retrospective study, prospective external validation in a multicenter setting is warranted before clinical implementation.
    Keywords:  age-related macular degeneration; arbitration; artificial intelligence; deep learning; diabetic retinopathy; double reading; fundus photography; retinal vein occlusion; screening
    DOI:  https://doi.org/10.3390/diagnostics16172707
  16. Front Med (Lausanne). 2026 ;13 1934432
       Background: Patients with diabetes increasingly consult artificial intelligence (AI) chatbots for medical advice, including guidance on antidiabetic medication management during Ramadan fasting, because AI can simplify and summarize long, complex guidelines. Also, in hospital settings, these tools are being used in hospitals much faster than it takes to establish formal regulations and guidelines for their use. Evaluations of the accuracy, completeness, and reproducibility of such advice across languages are still lacking. Therefore, the study aims to evaluate and compare the accuracy, completeness, safety, and reproducibility of three widely used AI chatbots-ChatGPT, Google Gemini, and Microsoft Copilot-when providing antidiabetic medication adjustment advice during Ramadan in both English and Arabic.
    Methods: Twenty-three standardized clinical scenarios covering common antidiabetic regimens were presented to each chatbot in both English and Arabic. Each query was repeated to evaluate reproducibility, resulting in 276 responses scored. Responses were assessed against the International Diabetes Federation-Diabetes and Ramadan (IDF-DAR) Guidelines using a 0-2 accuracy scale, a 0-4 completeness scale, and a 0-3 safety scale.
    Results: Overall, 77% of responses were fully consistent with the guideline, 12% were partially consistent, and 11% (30/276) contained clinically harmful or contradictory advice; harmful responses were about twice as common in Arabic as in English (14% vs. 8%). Completeness and safety were high, with medians at the observed ceiling. In the generalized linear mixed models, chatbots did not differ significantly in accuracy, completeness, or safety, and there was no significant main effect of language or chatbot × language interaction; the strongest signals were a chatbot effect on completeness (p = 0.068) and a language effect on safety (p = 0.064), both non-significant. Two-week reproducibility was fair for accuracy (weighted κ = 0.20, p = 0.009) and completeness (κ = 0.29, p = 0.001) and showed a very low κ in the safety scale (κ = 0.02, p = 0.81).
    Conclusions: AI chatbots demonstrated comparable performance in delivering guideline-based advice for diabetes management during Ramadan, with no significant differences in accuracy, completeness, or safety. While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time. These results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.
    Keywords:  ChatGPT; Copilot; Gemini; Ramadan; artificial intelligence; diabetes mellitus; medication safety
    DOI:  https://doi.org/10.3389/fmed.2026.1934432
  17. Ann Acad Med Singap. 2026 Sep 11.
      
    Keywords:  access to information; artificial intelligence; data analysis; data collection; diabetes mellitus; public health informatics
    DOI:  https://doi.org/10.47102/annals-acadmedsg.2026124
  18. Front Public Health. 2026 ;14 1942799
       Background: Patients with painful diabetic peripheral neuropathy (PDPN) increasingly use generative artificial intelligence chatbots for information on symptoms, treatment, foot care, and when to seek professional help. Their usefulness depends on safety, accuracy, guideline concordance, actionability, and readability.
    Objective: To compare five publicly accessible generative AI chatbots in answering standardized patient-oriented questions about PDPN. Methods: Sixty standardized English-language questions covering eight clinical domains were submitted once to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao in separate single-turn conversations, yielding 300 responses. Five reviewers independently assessed safety, accuracy, guideline concordance, and actionability using predefined criteria. Guideline concordance was scored against six mapped elements per question and converted to a percentage. Actionability was assessed using seven binary criteria with prespecified question-level applicability. Readability was evaluated using six established indices. Paired comparisons used Cochran's Q test for safety and Friedman tests for non-binary outcomes, followed by multiplicity-adjusted pairwise analyses.
    Results: All 300 responses were analyzed. Inter-rater agreement was high for safety (Fleiss' κ = 0.874), accuracy [ICC (2,1) = 0.881], guideline concordance [ICC (2,1) = 0.874], and actionability [ICC (2,1) = 0.877]. Twenty-eight responses (9.3%) were classified as unsafe or potentially unsafe. Unsafe-response rates ranged from 5.0 to 15.0%, with no detected overall between-model difference (Cochran's Q = 4.462, p = 0.347). Accuracy, guideline concordance, and actionability differed across models (all p < 0.001; Kendall's W = 0.683, 0.730, and 0.556, respectively). ChatGPT generally achieved higher content-related scores, whereas Doubao scored lower. All readability indices also differed across models (all p < 0.001), with ChatGPT and Doubao producing less complex text and DeepSeek showing greater reading difficulty.
    Conclusion: The five chatbots showed distinct performance patterns across content quality and readability. Although no overall safety difference was detected, every system generated at least one response with a plausible pathway to inappropriate self-management, delayed assessment, medication or product misuse, or preventable injury. Chatbots may support general patient education, but medication decisions, foot-risk assessment, and urgent-care triage require professional verification.
    Keywords:  actionability; chatbot; generative artificial intelligence; guideline concordance; painful diabetic peripheral neuropathy; readability; safety
    DOI:  https://doi.org/10.3389/fpubh.2026.1942799