IMPROVING THE SYSTEM FOR EXTRACTING, GATHERING, AND SUMMARIZING FLOOD EVENTS FROM ONLINE NEWS WITH LARGE LANGUAGE MODELS
Main Article Content
Abstract
This research aims to enhance the efficiency of a semi-automated system for extracting, gathering, and spatially analyzing flood event data from online news in Thailand. It addresses the limitations of the previous system, which was built on Python and Named Entity Recognition (NER) techniques like WangchanBERTa model, which suffered from complexity in handling unstructured data and multi-step workflows. The proposed system leverages a JavaScript and Node.js-based architecture to manage web scraping tasks, integrated with n8n workflow automation as the user interface, and utilizes Gemini Flash 2.5 API (a Large Language Model: LLM) to process natural language, extract spatiotemporal entities (date, location, and province), summarize news content, and evaluate disaster severity levels. The system was evaluated using a dataset of 555 flood-related news articles from 2024. The experimental results demonstrated that the proposed system reduced the total processing time from 1,125 seconds (19 minutes) to 788 seconds (13 minutes), indicating a 29.95% reduction in computation time. In terms of data quality, the LLM exhibited higher extraction accuracy than the traditional NER, successfully handling complex sentence structures, resolving linguistic ambiguities, and capturing additional dimensions such as severity ratings that the previous system could not classify. The findings suggest that the redesigned system offers superior development flexibility and cross-platform integration, facilitating high-quality data preparation for heatmaps and policy-level spatial analysis to support semi-real-time warning systems. Nevertheless, deploying this system at scale requires careful consideration of data privacy and security during external API calls, as well as the cost management of language model token consumption.
Article Details
References
ชวลิต โควีระวงศ์, รัตถชล อ่างมณี, ไพรินทร์ มีศรี และ ดาวรถา วีระพันธ์. (2568). การพัฒนาและประยุกต์ใช้เทคนิคเว็บสแครปปิ้งและการประมวลผลภาษาธรรมชาติสำหรับการรวบรวมและวิเคราะห์ข้อมูลข่าวสารเกี่ยวกับน้ำท่วมในประเทศไทย. ใน: การประชุมวิชาการระดับชาติและนานาชาติครั้งที่ 4 ด้านวิทยาศาสตร์และเทคโนโลยี. วันที่ 7 มีนาคม 2568. คณะวิทยาศาสตร์และเทคโนโลยี มหาวิทยาลัยราชภัฏเพชรบูรณ์, เพชรบูรณ์. 314-323.
ศูนย์ข้อมูลสาธารณภัย. (2567). สรุปสถิติสาธารณภัยประจำปี พ.ศ. 2567 ของกรมป้องกันและบรรเทาสาธารณภัย (ปภ.). สืบคนเมื่อวันที่ 8 สิงหาคม 2569, จาก https://datacenter.disaster.go.th/datacenter/cms/8670?id=132971.
Deelman, E., Vahi, K., Juve, G., Rynge, M., Callaghan, S., Maechling, P. J., Mayani, R., Chen, W., Ferreira da Silva, R., Livny, M., & Wenger K. (2015). Pegasus, a workflow management system for science automation. Future Generation Computer Systems, 46, 17-35. doi:10.1016/j.future.2014.10.008.
Eze, E., Feyisetan, O., & Omojola, A. (2020). Named entity recognition for disaster management: A review and framework. International Journal of Disaster Risk Reduction, 48, 101612. doi:10.1016/j.ijdrr.2020.101612.
Google Developers. (2024). Google AI Studio Pricing and Model Documentation. Google AI for Developers. Retrieved June 25, 2026 form https://ai.google.dev/pricing.
Google Gemini Team. (2023). Gemini: A family of highly capable multimodal models. Retrieved August 1, 2026 from https://arxiv.org/pdf/2312.11805.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Jin Bang, Y., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1-38. doi:10.1145/3571730.
Lei, K., Ma, Y., & Tan, Z. (2014). Performance comparison and evaluation of Web development technologies in PHP, Python, and Node.js. In: Proceedings of IEEE International Conference on Computational Science and Engineering, 19-21 December, 2014, NW Washington, DC United States. 661-668. doi:10.1109/CSE.2014.142.
Lowphansirikul, L., Polpanumas, C., Jantrakulchai, N., & Nutanong, S. (2021). WangchanBERTa: Pretraining transformer-based Thai Language Models. Retrieved July 12, 2026 from https://arxiv.org/pdf/2101.09635.
Makarov, I. S., Larin, D. V., Vorobeva, E. G., Emelin, D. P., & Kartashov, D. A. (2025). The impact of asynchronous and multithreaded query processing models on the performance of server-side web applications. Software Systems and Computational Methods, (1), 13–20. doi.org/10.7256/2454-0714.2025.1.73665.
Phatthiyaphaibun, W., Chaovavanich, K., Polpanumas, C., Suriyawongkul, A., Lowphansirikul, L., Chormai, P., Limkonchotiwat, P., Suntorntip, T., & Udomcharoenchaikit, C. (2023). PyThaiNLP: Thai Natural Language Processing in Python. In: Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software, December 6, 2023, 25-36. doi:10.18653/v1/2023.nlposs-1.4.
Revilla-Romero, B., Hirpa, F. A., Thielen-del Pozo, J., Salamon, P., Pappenberger, F., & De Groeve, T. (2015). On the use of global flood forecasts and satellite-derived inundation maps for flood monitoring in data-sparse regions. Remote Sensing, 7(11), 15702-15728. doi:10.3390/rs71115702.
Singrodia, V., Mitra, A., & Paul, S. (2019). A review on web scraping and its applications. In: International Conference on Computer Communication and Informatics (ICCCI), 23-25 January, 2019, Coimbatore, India. 1-6. doi:10.1109/ICCCI.2019.8821809.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998-6008.
Wei, X., Cui, X., Cheng, N., Wang, X., Zhang, X., Huang, S., Xie, P., Xu, J., Chen, Y., Zhang, M., Jiang, Y., & Han, W. (2023). ChatIE: Zero-Shot Information Extraction via Chatting with ChatGPT. Retrieved August 8, 2026 from https://arxiv.org/pdf/2302.10205.