Formal Verification of Ethical Behavior in Mental Health Conversational Systems Using Finite-State Models

Main Article Content

Tongjai Yampaka
duangjai noolek
Orawan Chunhapran

Abstract

The evaluation of mental health conversational systems has relied on benchmark-driven and performance-based metrics, which are limited in their ability to guarantee ethical safety. Although conversational systems achieve high performance scores, critical ethical failures may occur in contexts involving emotional vulnerability, medical requests, or potential selfharm. To address this limitation, this study proposes a formal verification framework based on finite-state machine (FSM) model checking to assess ethical behavior. The proposed framework models multi-turn conversations as nite execution traces and uses expert-driven risk taxonomies to define ethical safety requirements. Ethical requirements are applied as accepted transitions and forbidden transitions, enabling deterministic verification of conversational trajectories. The approach offers reproducible and platform-independent ethical assurance based on formal methods. The study evaluated multiple conversational AI platforms using fixed-length sessions and a balanced set of risk scenarios. The difference results across platforms show that ethically unsafe trajectories can still emerge despite controlled experimental conditions. The ethical compliance rate, risk-level violation rate, and forbidden transition frequency successfully expose behavioral failure modes that remain hidden under conventional evaluation paradigms. The study supports a shift from performance-oriented evaluation toward verification-driven ethical assurance. In addition, FSM-based verification is demonstrated in promise for safety-critical conversational AI systems.

Article Details

How to Cite
[1]
T. Yampaka, duangjai noolek, and O. Chunhapran, “Formal Verification of Ethical Behavior in Mental Health Conversational Systems Using Finite-State Models”, ECTI-CIT Transactions, vol. 20, no. 4, pp. 643–657, Sep. 2026.
Section
Research Article

References

A. A. Abd-Alrazaq, A. Rababeh, M. Alajlani, B. M. Bewick and M. Househ, “Effectiveness and Safety of Using Chatbots to Improve Mental Health: Systematic Review and Meta-Analysis,” Journal of Medical Internet Research, vol. 22, no.7, p. e16021, 2020.

L. Laranjo et al., “Conversational agents in healthcare: A systematic review,” Journal of the American Medical Informatics Association, vol. 25, no. 9, pp. 1248–1258, 2018.

A. S. Miner, A. Milstein, S. Schueller, R. Hegde, C. Mangurian and E. Linos, “Smartphone-based conversational agents and responses to questions about mental health, interpersonal violence, and physical health,” JAMA Internal Medicine, vol. 176, no. 5, pp. 619–625, 2016.

P. P. Liang, C. Wu, L.-P. Morency and R. Salakhutdinov, “Towards understanding and mitigating social biases in language models,” arXiv preprint arXiv:2106.13219, 2021.

Y. Hua et al., “Standardizing and Scaffolding Health Care AI-Chatbot Evaluation: Systematic Review,” JMIR AI, vol. 4, p. e69006, 2025.

Y. Bai et al., “Constitutional AI: Harmlessness from AI feedback,” arXiv preprint arXiv:2212.08073, 2022.

E. Perez et al., “Red teaming language models with language models,” in Proceedings of the Empirical Methods Natural Language Processing (EMNLP), pp. 3419–3448, 2022.

Z. Zhang et al., “SafetyBench: Evaluating the safety of large language models,” arXiv preprint arXiv:2309.07045, 2023.

S. Gehman, S. Gururangan, M. Sap, Y. Choi and N. A. Smith, “RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models,” arXiv preprint arXiv:2009.11462, 2020.

Y. Liu et al., “G-Eval: NLG evaluation using GPT-4 with better human alignment,” in Proceedings of the Empirical Methods Natural Language Processing (EMNLP), 2023.

L. Zheng et al., “Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,” in Proceedings of the 37th International Conference on Neural Information Processing Systems, no. 2020, pp. 46595-46623, 2023.

C. Baier and J.-P. Katoen, Principles of Model Checking. Cambridge, MA, USA: MIT Press, 2008.

L. Dennis, M. Fisher, M. Slavkovik and M. Webster, “Formal verification of ethical choices in autonomous systems,” Robotics and Autonomous Systems, vol. 77, pp. 1-14, 2016.

E. M. Clarke, W. Klieber, M. Nov´aˇcek and P. Zuliani, “Model checking and the state explosion problem,” Tools for Practical Software Verification, pp. 1–30, 2012.

H. Li, R. Zhang, Y.-C. Lee, R. E. Kraut and D. C. Mohr , “Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being,” npj Digital Medicine, vol. 6, no. 236, 2023.

M. Luckcuck, M. Farrell, L. A. Dennis, C. Dixon and M. Fisher, “Formal Specification and Verification of Autonomous Robotic Systems: A Survey,” ACM Computing Surveys, vol. 52, no. 5, pp. 1-41, 2019.

M. A. Kuhail et al., “A systematic review on mental health chatbots: Trends, design principles, evaluation methods and future research agenda,” Human Behavior and Emerging Technologies, no. 1, 2025.

E. L. Bunge and C. Desage, “A framework for evaluating mental health artificial intelligence-based conversational agents,” Journal of Technology in Behavioral Science, vol. 10, pp. 731739, 2025.

S. Bensalem et al., “Bridging formal methods and machine learning with model checking and global optimisation,” Journal of Logical and Algebraic Methods in Programming, vol. 137, p. 100941, 2024. 657

Y. Y. Elboher et al., “Formal verification of neural networks with early exits,” arXiv preprint arXiv:2512.20755, Dec. 2025.

D. Golpayegani, J. Hovsha, L. W. S. Rossmaier, R. Saniei and J. Miˇsi´c, “Towards a Taxonomy of AI Risks in the Health Domain,” 2022 Fourth International Conference on Transdisciplinary AI (TransAI), Laguna Hills, CA, USA, pp. 1-8, 2022.

Z. Porter, I. Habli, J. McDermid and M. Kaas , “A principles-based ethics assurance argument pattern for AI and autonomous systems,” AI and Ethics, vol. 4, pp. 593–616, 2024.

R. Bommasani, P. Liang and T. Lee, “Holistic evaluation of language models,” Annals of the New York Academy of Sciences, vol. 1525, no. 1, pp. 140–143, 2023.

M. Bdiwi, I. A. Naser, J. Halim, S. Bauer, P. Eichler and S. Ihlenfeldt “Towards Safety 4.0: A novel approach for flexible human-robot interaction based on safety-related dynamic finitestate machine with multilayer operation modes,” Frontiers Robotics and AI, vol. 9, 2022.

K. P. Jevitha, B. Jayaraman and M. Sethumadhavan, “Runtime verification on abstract finite state models,” Journal of Systems and Software, vol. 216, p. 112138, 2024.