DEAR EDITOR,
We thank the author(s) of the Letter to the Editor entitled “Letter to the Editor Re: Free Large Language Model-based Chatbots Can Help to Align Statistical Tests with the Study Design and Avoid Pseudo-replication” for their careful reading of our article, “AI in Patient Care: Evaluating Large Language Model Performance Against Evidence-Based Guidelines for Pulmonary Embolism,” and for highlighting an important statistical issue.1, 2
The main concern relates to the use of the Kruskal-Wallis test for comparing scores from four large language model (LLM)-based chatbots. We acknowledge that, because the same ten clinical questions were submitted to each model, the observations are better considered as blocked or repeated measures rather than independent groups. Therefore, a repeated-measures non-parametric approach is more appropriate for the overall comparison.3, 4
To address this point, we reanalyzed the model scores using the Friedman test, treating each question as a block and the four LLMs as related conditions. The Friedman test showed no statistically significant difference among the four models [χ2=4.50, P = 0.212]. Since the overall test was not significant, no post-hoc pairwise comparisons were performed.5 This result is consistent with the main conclusion of our original analysis: although ChatGPT-4o achieved the highest total score and DeepSeek-V2 the lowest, no model demonstrated statistically significant overall superiority.
We agree that the statistical analysis section could have more clearly reflected the repeated-measures structure of the data. Nevertheless, the reanalysis did not change the principal interpretation of the study. The evaluated LLMs demonstrated model-specific strengths and limitations across different components of pulmonary embolism management, but none consistently outperformed the others across all domains.
Accordingly, the relevant Methods and Results statements in the original article should be corrected through an erratum, replacing the Kruskal-Wallis analysis with the Friedman test and reporting the updated result [χ2(3)=4.50, P = 0.212].
We appreciate the opportunity to clarify this issue and believe that this discussion contributes constructively to the methodological rigor of studies evaluating AI-based clinical decision-support tools.


