← Back to all studies
anti-inflammatory foods

Comparing the Accuracy of ChatGPT-4o, DeepSeek-V3, and Gemini 2.5 Flash in Answering Frequently Asked Questions About Systemic Lupus Erythematosus: Quantitative Study

8 June 2026

Widhani A, Maria S, Chairani AP, Visco NYA, Nurhadi MFA, Harjoprawito LA, Elvira D, Tedja I, Yuniza Y, Fetarayani D, Koesnoe S, Wicaksana B, Hasibuan AS, Yunihastuti E.

Summary

What the study found

Researchers compared how well three popular AI tools answered common questions about Systemic Lupus Erythematosus (SLE), a complex autoimmune condition. While all tools provided generally satisfactory information, Gemini 2.5 Flash emerged as the most accurate among them, though its answers were also the most detailed.

Key findings

  • Gemini 2.5 Flash achieved the highest accuracy ratings from clinical immunologists, outperforming both DeepSeek-V3 and ChatGPT-4o.
  • There is a significant readability gap, as all three chatbots produced complex text that may be difficult for the average patient to understand.
  • A negative correlation was observed between word count and errors, meaning longer, more comprehensive responses tended to be more factually correct.
  • The absolute difference in accuracy between the three AI models was relatively small, suggesting all three are becoming increasingly reliable for general medical queries.

Practical takeaways

If you use AI to research anti-inflammatory diets or chronic disease management, prioritize tools like Gemini for technical accuracy but remain aware that the language may be overly clinical. Always cross-reference AI-generated health advice with a medical professional to ensure the information is translated correctly for your specific health needs.

Limitations

The study focused specifically on questions asked in Bahasa Indonesia, so the performance levels might vary for English or other languages. Additionally, the researchers did not evaluate the clinical safety or potential for harmful advice, only the factual accuracy and reading level.

Abstract

<h4>Background</h4>Systemic lupus erythematosus (SLE) is a complex, fluctuating disease, creating a continuous need for reliable patient information. A prior study concluded that patients with SLE often turn to the internet, including artificial intelligence (AI) chatbots, for information regarding SLE. The rise of AI chatbots as a primary information source presents a critical challenge regarding the accuracy of the information they provide.<h4>Objective</h4>This study aimed to evaluate the performance of the latest generation of AI chatbots (ChatGPT-4o, DeepSeek-V3, and Gemini 2.5 Flash) in answering frequently asked questions about SLE.<h4>Methods</h4>Twenty-two frequently asked questions about SLE in Bahasa Indonesia (the Indonesian language) were posed to each chatbot. Responses were independently and blindly evaluated for accuracy by 5 clinical immunologists using a 4-point Likert scale. Readability was assessed using the Flesch reading ease score formula. Statistical comparisons for accuracy and readability were performed using repeated-measures ANOVA or the Friedman test, followed by the Bonferroni test for pairwise comparisons. The Spearman ρ was used to evaluate correlations among accuracy, readability, and word count.<h4>Results</h4>Gemini 2.5 Flash demonstrated the highest accuracy, with a mean score of 1.25 (SD 0.53), significantly outperforming ChatGPT-4o (mean 1.71, SD 0.61; P<.001). Gemini 2.5 Flash significantly outperformed ChatGPT-4o in 2 evaluated domains. The interreliability analysis revealed a statistically significant level of agreement among the 5 evaluators across all responses (Kendall W=0.389; P<.001). Readability for all 3 chatbots was low (median Flesch reading ease score 42.22-46.66). Gemini 2.5 Flash produced the longest responses (8509 total words), followed by DeepSeek-V3 (5410 words) and ChatGPT-4o (3632 words). A significant negative correlation was found between word count and lower accuracy (ρ=-0.401; P=.001).<h4>Conclusions</h4>Our study found that ChatGPT-4o, DeepSeek-V3, and Gemini 2.5 Flash provided overall satisfactory responses to SLE-related questions. The highest accuracy was demonstrated by Gemini 2.5 Flash; however, the absolute differences in scores among the 3 AI chatbots were relatively small. All 3 AI chatbots demonstrated low readability, which may limit accessibility for patient use. This finding highlights a critical "blind spot" in which clinical accuracy, as rated by experts, does not equate to patient accessibility. Thus, further research is required to develop more comprehensive evaluation frameworks incorporating safety, factuality, and calibration of AI chatbots across different medical fields and topics.
Source study →