Comparison of LLMs for AI Chatbots in Local Government Public Services
Main Article Content
Abstract
This research aims to develop and compare the performance of AI chatbot systems integrated with three large language models: GPT-4o mini, Gemini 2.5 Flash, and Claude 3.5 Haiku. Installed on Line Official Accounts (LINE OA) via a webhook architecture, the systems utilized Google Apps Script as middleware. Chatbot performance was evaluated by 141 Village Health Volunteers (VHVs) from Ban Khok Subdistrict, Sangkhom District, Udon Thani Province, Thailand, selected through purposive sampling. The evaluation focused on six dimensions—accuracy, completeness, relevance, natural language, clarity, and reliability—under the same environment and knowledge base in Google Sheets. Chatbot responses were validated prior to testing with the sample group. Inferential statistical analysis using Repeated-Measures ANOVA revealed statistically significant differences in response speed, F(2, 98) = 37.941, p < .001, with GPT-4o mini demonstrating the fastest average speed of 7.09 seconds. In terms of user satisfaction, Claude 3.5 Haiku scored the highest in natural language (92.20%) and reliability (91.40%), which were statistically significant (p < .001), as well as in completeness (p = .003). Meanwhile, GPT-4o mini scored highest in relevance at 92.40% (p < .001). These results indicate that no single model was completely superior across all evaluated dimensions.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
All authors need to complete copyright transfer to Journal of Applied Informatics and Technology prior to publication. For more details click this link: https://ph01.tci-thaijo.org/index.php/jait/copyrightlicense