Front Public Health. 2026 ;14
1942799
Background: Patients with painful diabetic peripheral neuropathy (PDPN) increasingly use generative artificial intelligence chatbots for information on symptoms, treatment, foot care, and when to seek professional help. Their usefulness depends on safety, accuracy, guideline concordance, actionability, and readability.
Objective: To compare five publicly accessible generative AI chatbots in answering standardized patient-oriented questions about PDPN. Methods: Sixty standardized English-language questions covering eight clinical domains were submitted once to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao in separate single-turn conversations, yielding 300 responses. Five reviewers independently assessed safety, accuracy, guideline concordance, and actionability using predefined criteria. Guideline concordance was scored against six mapped elements per question and converted to a percentage. Actionability was assessed using seven binary criteria with prespecified question-level applicability. Readability was evaluated using six established indices. Paired comparisons used Cochran's Q test for safety and Friedman tests for non-binary outcomes, followed by multiplicity-adjusted pairwise analyses.
Results: All 300 responses were analyzed. Inter-rater agreement was high for safety (Fleiss' κ = 0.874), accuracy [ICC (2,1) = 0.881], guideline concordance [ICC (2,1) = 0.874], and actionability [ICC (2,1) = 0.877]. Twenty-eight responses (9.3%) were classified as unsafe or potentially unsafe. Unsafe-response rates ranged from 5.0 to 15.0%, with no detected overall between-model difference (Cochran's Q = 4.462, p = 0.347). Accuracy, guideline concordance, and actionability differed across models (all p < 0.001; Kendall's W = 0.683, 0.730, and 0.556, respectively). ChatGPT generally achieved higher content-related scores, whereas Doubao scored lower. All readability indices also differed across models (all p < 0.001), with ChatGPT and Doubao producing less complex text and DeepSeek showing greater reading difficulty.
Conclusion: The five chatbots showed distinct performance patterns across content quality and readability. Although no overall safety difference was detected, every system generated at least one response with a plausible pathway to inappropriate self-management, delayed assessment, medication or product misuse, or preventable injury. Chatbots may support general patient education, but medication decisions, foot-risk assessment, and urgent-care triage require professional verification.
Keywords: actionability; chatbot; generative artificial intelligence; guideline concordance; painful diabetic peripheral neuropathy; readability; safety