An Unsupervised Machine Learning Model for Identifying Suicidal Posts Using Latent Dirichlet Allocation and Text Feature Extraction

Main Article Content

Natthaphong Suthamno
Jessada Tanthanuch

Abstract

This study aims to enhance the performance of unsupervised machine learning models in identifying suicidal behavior from English-language Reddit posts. Using Python, a model was developed and trained on 232,074 posts with Latent Dirichlet Allocation (LDA) for topic modeling. Text feature extraction techniques incorporating unigram to 4-gram models were applied prior to model input. Four experimental setups tested different feature word sizes: the top 50,000, 100,000, 150,000 most frequent words, and all available features. Results showed that using the top 50,000 features yielded the highest accuracy (85.45%) but the lowest recall (95.55%), while using all features achieved lower accuracy (79.54%) but higher recall (97.70%). The findings suggest that reducing feature quantity enhances model accuracy and focus on suicide-related content, but may compromise recall due to loss of contextual information. Overall, the study concludes that LDA-based models can perform effectively without relying on the full vocabulary of large datasets, though a trade-off between accuracy and recall must be considered.

Article Details

Section
Research Paper

References

World Health Organization (2023), Depressive disorder (depression). Available Online at https://www.who.int/news-room/fact-sheets/detail/depression, accessed on 19 April 2025.

A. Clark, C. Fox, and S. Lappin. The Handbook of Computational Linguistics and Natural Language Processing. Hoboken, NJ, USA: John Wiley & Sons, 2013.

P. Burnap, W. Colombo, and J. Scourfield. "Machine classification and analysis of suicide-related communication on Twitter." 26th ACM Conference on Hypertext & Social Media (HT '15), New York, NY, USA, pp. 75-84, 2015.

G. Astoveza, R. J. P. Obias, R. J. L. Palcon, R. L. Rodriguez, B. S. Fabito, and M. V. Octaviano. "Suicidal behavior detection on Twitter using neural network." TENCON 2018 - 2018 IEEE Region 10 Conference, Jeju, Korea (South), pp. 0657-0662, 2018.

S. Jain, S. P. Narayan, R. K. Dewang, U. Bhartiya, N. Meena, and V. Kumar. "A Machine Learning based Depression Analysis and Suicidal Ideation Detection System using Questionnaires and Twitter." 2019 IEEE Students Conference on Engineering and Systems (SCES), Allahabad, India, pp. 1-6, 2019.

M. M. Tadesse, H. Lin, B. Xu, and L. Yang. "Detection of Suicide Ideation in Social Media Forums Using Deep Learning." Algorithms, Vol. 13, No. 1, 2019.

F. M. Shah, F. Haque, R. U. Nur, S. Al Jahan, and Z. Mamud. "A Hybridized Feature Extraction Approach To Suicidal Ideation Detection From Social Media Post." 2020 IEEE Region 10 Symposium (TENSYMP), Dhaka, Bangladesh, pp. 985-988, 2020.

A. Chadha and B. Kaushik. "A Hybrid Deep Learning Model Using Grid Search and Cross-Validation for Effective Classification and Prediction of Suicidal Ideation from Social Network Data." New Generation Computing, Vol. 40, No. 4, pp. 889-914, 2022.

A. Abdulsalam and A. Alhothali. "Suicidal ideation detection on social media: A review of machine learning methods." Social Network Analysis and Mining, Vol. 14, No. 1, 2024.

S. Bird, E. Klein, and E. Loper. Natural Language Processing with Python. Sebastopol, CA, USA: O'Reilly Media, 2009.

N. Jones, N. Jaques, P. Pataranutaporn, A. Ghandeharioun, and R. Picard. "Analysis of Online Suicide Risk with Document Embeddings and Latent Dirichlet Allocation." 2019 8th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), Cambridge, MA, USA, pp. 1-5, 2019. doi: 10.1109/ACIIW.2019.8925047.

R. Kushwaha and P. Kaur. "A federated LDA-based approach for topic modelling." Intelligent Decision Technologies, Vol. 19, No. 4, pp. 2738-2759, 2025. doi: 10.1177/18724981251319629.

K. W. Ng, G. L. Tian, and M. L. Tang. Dirichlet and Related Distributions: Theory, Methods and Applications. Hoboken, NJ, USA: John Wiley & Sons, 2011.

B. A. Frigyik, A. Kapila, and M. R. Gupta. "Introduction to the Dirichlet Distribution and Related Processes." UWEE Technical Report Series, Vol. 1, pp. 1-27, 2010.

D. M. Blei, A. Y. Ng, and M. I. Jordan. "Latent Dirichlet Allocation." Journal of Machine Learning Research, Vol. 3, pp. 993-1022, 2003.

W. M. Darling. "A Theoretical and Practical Implementation Tutorial on Topic Modeling and Gibbs Sampling." 49th annual meeting of the association for computational linguistics: Human language technologies, pp. 642-647, 2011.

M. T. Pilehvar and J. Camacho-Collados. Embeddings in Natural Language Processing: Theory and Advances in Vector Representations of Meaning. San Rafael, CA, USA: Morgan & Claypool Publishers, 2020.

A. H. Razavi, D. Inkpen, D. Brusilovsky, and L. Bogouslavski. "General topic annotation in social networks: A Latent Dirichlet Allocation approach." in Advances in Artificial Intelligence: 26th Canadian Conf. Artificial Intelligence (Canadian AI 2013), Regina, SK, Canada, pp. 293-300, 2013.