Pradnya D. Bormane,
Mitali Balkrishna Satpute,
Onkar Chandrakant Rangate,
Rajratna Kadu Salve,
- Assistant Professor, Department of Artificial Intelligence and Data Science, AISSMS Institute of Information Technology, Maharashtra, India
- Student, Department of Artificial Intelligence and Data Science, AISSMS Institute of Information Technology, Maharashtra, India
- Student, Department of Artificial Intelligence and Data Science, AISSMS Institute of Information Technology, Maharashtra, India
- Student, Department of Artificial Intelligence and Data Science, AISSMS Institute of Information Technology, Maharashtra, India
Abstract
Public safety in urban settings, as these areas are crowded and experience increasing incidents of harassment, verbal abuse, and physical violence. Traditional CCTV surveillance systems heavily rely on human operators who have to continuously monitor the video feeds, which is inefficient and prone to human error. AI-based harassment and violence detection systems automatically monitor public environments in real-time using video and audio analysis. The system is based on a multilayer guardrail architecture that includes the use of several artificial intelligence models to enhance the reliability of the detection. The proposed system identifies three major threats: physical violence, verbal abuse, and inappropriate physical contact in crowded public places. Video analysis is done with the SlowFast R50 deep learning model; pose-based motion analysis is implemented with MediaPipe with velocity tracking, and speech toxicity detection is done with a multilingual DistilBERT classifier. Additionally, object detection using YOLOv12 is used to analyze proximity between individuals to determine suspicious physical contact. The system combines these layers of detection through a fusion engine that computes a final risk score and triggers alerts in cases where the risk is above a certain threshold. When an incident is detected, the system captures a video clip, takes an incident report, and determines the nearest police station to respond to an emergency. Experimental evaluation is found to have a good performance, with 97% accuracy for violence detection and a 92% F1 score for multilingual toxicity detection. The system shows promise for AI-driven surveillance for public safety improvement and lessening of reliance on manual surveillance.
Keywords: Action recognition, artificial intelligence, computer vision, deep learning, harassment detection, object detection, violence detection, video surveillance
[This article belongs to Journal of Advancements in Robotics ]
References
- Sultani K, Chen C, Shah M. Real-world anomaly detection in surveillance videos. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2018. p. 6479-88.
- Szydłowski M, et al. UCF-Crime: a benchmark for large-scale surveillance video anomaly detection [Preprint]. 2018. arXiv:1801.02160.
- Zhu Y, Newsam S. XD-Violence: a large-scale dataset for real-world violent video detection [Preprint]. 2020. arXiv:2003.05074.
- Lin J, Gan C, Han S. TSM: temporal shift module for efficient video understanding. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 2019. p. 7083-93.
- Carreira J, Zisserman A. Quo vadis, action recognition? A new model and the Kinetics dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017. p. 4724-33.
- Tran D, et al. Learning spatiotemporal features with 3D convolutional networks. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 2015. p. 4489-97.
- Feichtenhofer C, Fan H, Malik J, He K. SlowFast networks for video recognition. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 2019. p. 6202-11.
- Bertasius G, Wang H, Torresani L. Is space-time attention all you need for video understanding? In: Proceedings of the International Conference on Machine Learning (ICML); 2021.
- Arnab A, et al. ViViT: a video vision transformer. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 2021. p. 6836-46.
- Redmon J, Divvala S, Girshick R, Farhadi A. You only look once: unified, real-time object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016. p. 779-88.
- Jocher G. YOLOv5 by Ultralytics [Internet]. GitHub; 2020. Available from: https://github.com/ultralytics/yolov5
- Wang CY, Bochkovskiy A, Liao HYM. YOLOv7: trainable bag-of-freebies [Preprint]. 2022. arXiv:2207.02696.
- Dosovitskiy A, et al. An image is worth 16×16 words: transformers for image recognition at scale. In: International Conference on Learning Representations (ICLR); 2021.
- Sudhakaran S, Lanz O. Learning to detect violent videos using convolutional neural networks. In: IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS); 2017. p. 1-6.
- Hassner M, Itcher Y, Kliper-Gross O. Violent flows: real-time detection of violent crowd behavior. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPR Workshops); 2012. p. 1-6.
- Clarin C, Dionisio R, Castro J. Real-time detection of violent actions in surveillance videos using deep learning. In: IEEE Region 10 Conference (TENCON); 2019. p. 2638-43.
- Ionescu RT, et al. Object-centric auto-encoders for anomaly detection in video. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2019.
- Liu W, et al. Social network surveillance for abnormal behavior detection. Pattern Recognit. 2018;77:114-26.
- Hasan M, et al. Learning temporal regularity in video sequences. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016.
- Chong YS, Tay YH. Abnormal event detection in videos using spatiotemporal autoencoder [Preprint]. 2017. arXiv:1701.01546.
- Chong YS, Tay YH. Abnormal event detection in videos using GANs. In: Proceedings of the IEEE International Conference on Image Processing (ICIP); 2017.
- Medel JR, Savakis A. Anomaly detection in video using predictive convolutional long short-term memory networks [Preprint]. 2016. arXiv.
- Feng J, et al. Deep learning for anomaly detection in video surveillance: a survey. ACM Comput Surv. 2021;54(2):1-36.
- Liu Y, Ma C, et al. Spatiotemporal graph neural networks for anomaly detection. IEEE Trans Circuits Syst Video Technol. 2022.
- Liu Z, et al. Swin Transformer: hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 2021.
- Guo J, et al. Lightweight deep learning models for real-time surveillance. IEEE Access. 2023.
- Ren S, He K, Girshick R, Sun J. Faster R-CNN: towards real-time object detection. In: Advances in Neural Information Processing Systems (NeurIPS); 2015.
- Lin TY, et al. Microsoft COCO: common objects in context. In: European Conference on Computer Vision (ECCV); 2014.
- Feichtenhofer C, et al. Spatiotemporal feature learning for action recognition. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 2017.
- Zhan B, et al. Deep learning-based violence detection: a review. Multimed Tools Appl. 2022.

Journal of Advancements in Robotics
| Volume | 13 | |
| Issue | 02 | |
| Received | 03/04/2026 | |
| Accepted | 27/06/2026 | |
| Published | 29/06/2026 | |
| Publication Time | 87 Days |