Robust Violence Detection in Heterogeneous Surveillance Videos using a Dual-Branch Spatiotemporal Network

Main Article Content

Sai Thu Ya Aung
Worapan Kusakunniran
Yu Nandar Aung
Chattapatr Leeraha

Abstract

Automatic detection of violent events in surveillance video is a critical component of modern public safety systems. In contrast, existing methods often show limited generalization, performing well on individual datasets but degrading under domain shifts in real-world environments. To address this generalization challenge, this paper proposes an Asymmetric Dual- Branch Network trained on a large-scale Multi-Source Dataset integrating nine heterogeneous benchmark datasets. The proposed architecture combines a motion branch based on ResNet-18 trained from scratch to capture motion-specific representations with a spatial branch employing an X3D backbone selected through comparative evaluation of four state-of-the-art architectures, to learn complementary spatial semantic representations. This asymmetric design effectively balances detection accuracy and real-time performance while improving robustness across heterogeneous scenarios. Extensive experiments using 5-fold cross-validation demonstrate stable performance across multiple datasets. Notably, the proposed framework achieves state-of-the-art results on the Bus-violence dataset (89.07%) and RLVS dataset (96.65%), while maintaining a high processing rate of 76.61 FPS. These empirical results, including validated performance on the unseen Kranok-NV dataset, confirm that the asymmetric integration of decoupled features offers a practical and scalable solution for real-time urban surveillance in uncontrolled environments.

Article Details

How to Cite
[1]
S. T. Y. . Aung, W. Kusakunniran, Y. N. Aung, and C. Leeraha, “Robust Violence Detection in Heterogeneous Surveillance Videos using a Dual-Branch Spatiotemporal Network”, ECTI-CIT Transactions, vol. 20, no. 3, pp. 562–574, Jul. 2026.
Section
Research Article

References

B. K. P. Horn and B. G. Schunck, “Determining optical flow,” Artificial Intelligence, vol. 17, no. 1–3, pp. 185–203, 1981.

N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), San Diego, CA, USA, vol. 1, pp. 886-893, 2005.

K. Yun, H. Jeong, K. M. Yi, S. W. Kim and J. Y. Choi, “Motion Interaction Field for Accident Detection in Traffic Surveillance Video,” 2014 22nd International Conference on Pattern Recognition, Stockholm, Sweden, pp. 3062-3067, 2014.

P. Khanarsa and S. Kitsiranuwat, “Deep learning-based ensemble approach for conventional pap smear image classification,” ECTI Transactions on Computer and Information Technology (ECTI-CIT), vol. 18, no. 1, pp. 101– 111, 2024.

S. Sudhakaran and O. Lanz, “Learning to detect violent videos using convolutional long shortterm memory,” 2017 14th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Lecce, Italy, pp. 1-6, 2017.

D. Sriveni and R. Loganathan, “Multiconstraints active learning assisted deepensemble spatio-textural feature learning model for violence detection in surveillance dataset,” Connection Science, vol. 37, no. 1, p. 2544539, 2025.

K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in Proceedings of the 28th International Conference on Neural Information Processing Systems(NeurIPS), vol. 1, pp. 568–576, 2014.

C. Feichtenhofer, A. Pinz and A. Zisserman, “Convolutional Two-Stream Network Fusion for Video Action Recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, pp. 19331941, 2016.

C. Feichtenhofer, H. Fan, J. Malik and K. He, “SlowFast Networks for Video Recognition,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), pp. 6201-6210, 2019.

S. Mekruksavanich and A. Jitpattanakul, “FallNeXt: A deep residual model based on multibranch aggregation for sensor-based fall detection,” ECTI Transactions on Computer and Information Technology (ECTI-CIT), vol. 16, no. 4, pp. 352–364, 2022.

N. Alabid, “Interpretation of spatial relationships by objects tracking in a complex streaming video,” ECTI Transactions on Computer and Information Technology (ECTI-CIT), vol. 15, no. 2, pp. 245–257, 2021.

S. Thu Ya Aung, W. Kusakunniran and Y. Kwang Hooi, “Leveraging Fusion Methods of Human Pose and Motion Dynamics for Accurate Violence Detection in Video Surveillance,” in IEEE Access, vol. 13, pp. 168116-168125, 2025

K. He, X. Zhang, S. Ren and J. Sun, “Deep Residual Learning for Image Recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, pp. 770-778, 2016.

C. Feichtenhofer, “X3D: Expanding Architectures for Efficient Video Recognition,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, pp. 200-210, 2020.

G. Bertasius, H. Wang and L. Torresani, “Is space-time attention all you need for video understanding?” in Proceedings of the 38th International Conference on Machine Learning (ICML), pp. 813–824, 2021.

Z. Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, pp. 9992-10002, 2021.

A. Howard et al., “Searching for MobileNetV3,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), pp. 1314-1324, 2019.

W. Kay et al., “The kinetics human action video dataset,” arXiv preprint arXiv:1705.06950, 2017.

J. Deng, W. Dong, R. Socher, L. -J. Li, Kai Li and Li Fei-Fei, “ImageNet: A large-scale hierarchical image database,” 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, pp. 248-255, 2009.

R. Goyal et al., “The “Something Something” Video Database for Learning and Evaluating Visual Common Sense,” 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, pp. 5843-5851, 2017.

M. M. Soliman, M. H. Kamal, M. A. El-Massih Nashed, Y. M. Mostafa, B. S. Chawky and D. Khattab, “Violence Recognition from Videos using Deep Learning Techniques,” 2019 Ninth International Conference on Intelligent Computing and Information Systems (ICICIS), Cairo, Egypt, pp. 80-85, 2019.

T. Hassner, Y. Itcher and O. Kliper-Gross, “Violent flows: Real-time detection of violent crowd behavior,” 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, Providence, RI, USA, pp. 1-6, 2012.

M. Cheng, K. Cai and M. Li, “RWF-2000: An Open Large Scale Video Database for Violence Detection,” 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, pp. 4183-4190, 2021.

S¸. Aktı, G. A. Tataro˘glu and H. K. Ekenel, “Vision-based Fight Detection from Surveillance Cameras,” 2019 Ninth International Conference on Image Processing Theory, Tools and Applications (IPTA), Istanbul, Turkey, pp. 1-6, 2019.

H. A. Huillcen Baca, F. L. Palomino Valdivia and J. C. Gutierrez Caceres, “Efficient human violence recognition for surveillance in real time,” Sensors, vol. 24, no. 2, p. 668, 2024.

L. Ciampi et al., “Bus violence: An open benchmark for video violence detection on public transport,” Sensors, vol. 22, no. 21, p. 8345, 2022.

M. B. de Paula, D. H. Salvadeo and D. M. de Araujo, “CamNuVem: A robbery dataset for video anomaly detection,” Sensors, vol. 22, no. 24, p. 10016, 2022.

M. Kulkarni and R. Chakraborty, “Violent Activity Detection on Public Transportation using Surveillance Footage,” 2023 16th International Conference on Developments in eSystems Engineering (DeSE), Istanbul, Turkiye, pp. 737-742, 2023.

L. Ciampi, C. Santiago, F. Falchi, C. Gennaro, and G. Amato, “In the wild video violence detection: An unsupervised domain adaptation approach,” SN Computer Science, vol. 5, no. 7, p. 834, 2024.

A. Kulkarni, S. Modi, O. Vyawahare, A. Dubbewar and N. L. Pariyal, “Real-time quarrel detection in surveillance videos,” International 573 Journal of Novel Research and Development (IJNRD), vol. 10, no. 12, Dec. 2025. ¨

A. Akba¸s, H. U¸cg¨un and A. A. Ali, “A Hybrid Deep Learning Approach for Violence Detection in Videos,” 2025 Innovations in Intelligent Systems and Applications Conference (ASYU), Bursa, Turkiye, pp. 1-6, 2025.

D. C. Senadeera, X. Yang, D. Kollias and G. Slabaugh, “CUE-Net: Violence Detection Video Analytics with Spatial Cropping, Enhanced UniformerV2 and Modified Efficient Additive Attention,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, pp. 48884897, 2024.

S. Jung, T. Song, Y. Lee and S. Lee, “Shortwindow sliding learning for real-time violence detection via LLM-based auto-labeling,” arXiv preprint arXiv:2511.10866, 2025.

J. Chen, “Method and analysis of violent behavior recognition based on multimodal information fusion,” Highlights in Science, Engineering and Technology, vol. 161, pp. 34–42, 2026.

H. Huang and Q. Jiang, “IDG-ViolenceNet: A video violence detection model integrating identity-aware graphs and 3D-CNN,” Sensors, vol. 25, no. 20, p. 6272, 2025.

A. Pandey and P. Kumar, “BGRU-MTRA: Bilinear GRU networks with multi-path temporal residual attention for suspicious activity recognition,” Neural Computing and Applications, vol. 37, no. 1, pp. 185–212, 2025.

F. Meng, L. Zou, J. Lin and Z. Liu, “HSTNet: Violent action detection,” Applied Sciences, vol. 16, no. 4, p. 1825, 2026.

M. Mahmoud, B. Yagoub, M. F. Senussi, M. Abdalla, M. S. Kasem and H.-S. Kang, “Two-stage video violence detection framework using GMFlow and CBAM-enhanced ResNet3D,” Mathematics, vol. 13, no. 8, p. 1226, 2025.

D. C. Senadeera, X. Yang, S. Li, M. Awais, D. Kollias, and G. Slabaugh, “Dual branch VideoMamba with gated class token fusion for violence detection,” arXiv preprint arXiv:2506.03162, 2025.

E. Veltmeijer, M. Franken and C. Gerritsen, “Real-time violence detection and localization through subgroup analysis,” Multimedia Tools Appl., vol. 84, no. 7, pp. 3793–3807, 2025.

A. Kavathia and S. Sayer, “Optimizing violence detection in video classification accuracy through 3D convolutional neural networks,” arXiv preprint arXiv:2411.01348, 2024.

K. B. Kwan-Loo, J. C. Ort´ız-Bayliss, S. E. Conant-Pablos, H. Terashima-Mar´ın and P. Rad, “Detection of Violent Behavior Using Neural Networks and Pose Estimation,” in IEEE Access, vol. 10, pp. 86339-86352, 2022.