Cureus. 2026 Aug;18(8):
e115164
Background The growing adoption of artificial intelligence (AI) technologies, including large language models (LLMs), such as ChatGPT (OpenAI, San Francisco, CA) and many others, has expanded their role in statistical analysis and scientific research workflows. Despite this increasing use, conventional statistical software, such as IBM SPSS (IBM Corp., Armonk, NY), remains the accepted reference standard because of its established methodological precision, transparency, and reproducibility. Consequently, it is important to determine whether statistical results generated by LLMs achieve comparable levels of accuracy and consistency. As these systems undergo continual refinement, independent replication, comparative validation, and continuation studies are necessary to assess whether newer model generations produce findings that reliably correspond with those obtained using established statistical methodologies. Methodology Thirteen statistical procedures commonly used in clinical, medical, epidemiological, and applied health sciences research were evaluated using authentic datasets derived from previously published peer-reviewed studies and undergraduate applied health sciences coursework. Analyses included epidemiological tests, nonparametric analyses, chi-square analyses, univariable and multivariable logistic regression, Kaplan-Meier survival analysis, Cox proportional hazards regression, and measures of nominal association. For each procedure, dataset and variable names were copied directly from SPSS 31.0 data files and analyzed in GPT-5.5 using standardized prompts. AI-generated results were then compared directly with SPSS 31.0 outputs to evaluate computational agreement, methodological consistency, and inferential concordance across all statistical procedures. Results GPT-5.5 demonstrated a high degree of agreement with SPSS 31.0 across all statistical procedures evaluated, including risk and odds ratios, Wilcoxon signed-rank, Mann-Whitney U, Kruskal-Wallis, Friedman's test, one-way and two-way chi-square, binary and multivariable logistic regression, Kaplan-Meier survival analysis, Cox proportional hazards model, and phi and Cramer's V. Most analyses produced identical test statistics, effect estimates, confidence intervals, and inferential conclusions. Minor numerical differences were observed only in selected approximation statistics, including standardized Z values, Wald statistics, and Cox regression parameter estimates, and were attributable to expected implementation differences in optimization routines, convergence criteria, tie corrections, or reporting conventions. Importantly, none of these differences altered statistical significance, effect interpretation, or the overall scientific conclusions. Conclusions GPT-5.5 demonstrated excellent statistical consistency with SPSS 31.0 across a broad range of statistical analyses, producing complete or near-complete agreement in nearly all computational and inferential outcomes. Minor numerical differences were limited to implementation-specific approximations and did not affect statistical significance or substantive interpretation. These findings support the use of GPT-5.5 as a reliable adjunct for statistical verification, interpretation, and research support while reinforcing that independent methodological oversight and validation with established statistical software remain essential for scientific research.
Keywords: ai-assisted statistics; alignment; chatgpt; ibm spss; large language models