Purpose: The popularity of artificial intelligence-based chatbots in healthcare is increasing rapidly. In dentistry, they are used for various purposes, including patient education, diagnosis, and treatment planning. This study aimed to evaluate the accuracy and quality of three different AI chatbots, including ChatGPT-4, DeepSeek R1, and Gemini, in answering frequently asked patient questions.
Methods: 25 questions, which were frequently asked by patients in the subject of periapical surgery, were posed to ChatGPT-4, DeepSeek R1, and Gemini. Two independent raters recorded and assessed the recorded responses using a 5-point Likert scale. Also, the quality of the answers was scored using the EQIP scale. ANOVA and Kruskal-Wallis tests were used for the variables. The Spearman correlation test was used to determine the correlation between Likert and EQIP scores.
Results: Geminis’ Likert score was significantly lower than ChatGPT-4 and DeepSeek R1 (p<0.001). There was no significant difference between ChatGPT-4 and DeepSeek R1 (p=0.48). The difference between quality scores was not statistically significant. EQIP and Likert scores within groups showed poor correlation.
Conclusion: Artificial intelligence-based large language models can provide reliable information to patients about periapical surgery. However, there are differences in accuracy and quality between different platforms.
Keywords: Artificial intelligence, dentistry, large language models, oral surgery