Home / Current Issue / Paper 1717084
Efficient Net-B3 with Convolutional Block Attention and Focal Loss for Road Accident Detection in CCTV Surveillance Videos
Subject area: Science,Engineering and Technology · Area of research: Computer Vision, Deep Learning
DOI: https://doi.org/10.64388/IREV9I10-1717084
Abstract
Automated road accident detection from CCTV footage is a critical public safety challenge in smart city environments. Prior work using finetuned AlexNet on the ckay16 dataset achieved only 68.0% accuracy, a true positive rate (TPR) of 77.4%, and a critically high false positive rate (FPR) of 42.6% — rendering the system impractical for live deployment. In this paper we present a fully algorithmic framework that addresses these shortcomings without collecting additional data. Our method replaces AlexNet with an EfficientNet-B3 backbone, attaches a Convolutional Block Attention Module (CBAM) to the final feature block, and trains using Focal Loss with label smoothing, a Weighted Random Sampler, and a cosine-annealing learning-rate schedule with warm restarts. At inference, five-crop Test-Time Augmentation (TTA) further reduces false positives. On the held-out ckay16 test set the system achieves 96.0% accuracy, 95.7% TPR, 3.8% FPR, macro-F1 of 96.0%, and ROC-AUC of 0.981 — a 28-point accuracy gain and 39-point FPR reduction over the AlexNet baseline, demonstrating significant improvement over existing baselines.
Keywords
Road accident detection, EfficientNet, CBAM, Attention mechanism, Focal loss, Test-time augmentation, Transfer learning, CCTV surveillance
References
[1] A. Zahid, T. Qasim, N. Bhatti, and M. Zia, "A data-driven approach for road accident detection in surveillance videos," Multimedia Tools and Applications, vol. 83, pp. 17217–17231, 2024.
[2] W. Sultani, C. Chen, and M. Shah, "Real-world anomaly detection in surveillance videos," in Proc. IEEE CVPR, 2018, pp. 6479–6488.
[3] ckay16, "Accident Detection From CCTV Footage," Kaggle, 2020. [Online]. Available: https://www.kaggle.com/datasets/ckay16/accident-detection-from-cctv-footage
[4] Y.-K. Ki and D.-Y. Lee, "A traffic accident recording and reporting model at intersections," IEEE Trans. Intell. Transp. Syst., vol. 8, no. 2, pp. 188–194, 2007.
[5] N. Rasheed, S. A. Khan, and A. Khalid, "Tracking and abnormal behavior detection in video surveillance," in Proc. WAINA, 2014, pp. 61–66.
[6] H. Tan, Y. Zhai, Y. Liu, and M. Zhang, "Fast anomaly detection in traffic surveillance video based on robust sparse optical flow," in Proc. ICASSP, 2016, pp. 1976–1980.
[7] W. Sultani, C. Chen, and M. Shah, "Real-world anomaly detection (MIL)," IEEE TPAMI, vol. 43, no. 11, pp. 3893–3908, 2021.
[8] D. Singh and C. K. Mohan, "Deep spatio-temporal representation for detection of road accidents," IEEE Trans. Intell. Transp. Syst., vol. 20, no. 3, pp. 879–887, 2019.
[9] W. Ullah et al., "CNN features with bi-directional LSTM for real-time anomaly detection," Multimedia Tools and Applications, vol. 80, no. 11, pp. 16979–16995, 2021.
[10] M. Tan and Q. V. Le, "EfficientNet: Rethinking model scaling for convolutional neural networks," in Proc. ICML, 2019, pp. 6105–6114.
[11] S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, "CBAM: Convolutional block attention module," in Proc. ECCV, 2018, pp. 3–19.
[12] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, "Focal loss for dense object detection," in Proc. IEEE ICCV, 2017, pp. 2980–2988.
[13] S. Mukherjee, S. Bhowmik, and R. Chatterjee, "Focal loss vs cross-entropy for imbalanced surveillance classification," Pattern Recognition Letters, vol. 167, pp. 122–129, 2023.
[14] J. Hu, L. Shen, and G. Sun, "Squeeze-and-Excitation Networks," in Proc. IEEE CVPR, 2018, pp. 7132–7141.
[15] N. Nasaruddin, K. Muchtar, A. Afdhal, and A. P. J. Dwiyantoro, "Deep anomaly detection through visual attention in surveillance videos," Journal of Big Data, vol. 7, no. 1, pp. 1–17, 2020.
How to cite this paper
@article{1717084,
author = {M Koteswara Rao, G Raja Sekhar Reddy, P Swarna Kamal, P S C S V Sainadh},
title = {Efficient Net-B3 with Convolutional Block Attention and Focal Loss for Road Accident Detection in CCTV Surveillance Videos},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {10},
pages = {3725-3731},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1717084.pdf},
abstract = {Automated road accident detection from CCTV footage is a critical public safety challenge in smart city environments. Prior work using finetuned AlexNet on the ckay16 dataset achieved only 68.0% accuracy, a true positive rate (TPR) of 77.4%, and a critically high false positive rate (FPR) of 42.6% — rendering the system impractical for live deployment. In this paper we present a fully algorithmic framework that addresses these shortcomings without collecting additional data. Our method replaces AlexNet with an EfficientNet-B3 backbone, attaches a Convolutional Block Attention Module (CBAM) to the final feature block, and trains using Focal Loss with label smoothing, a Weighted Random Sampler, and a cosine-annealing learning-rate schedule with warm restarts. At inference, five-crop Test-Time Augmentation (TTA) further reduces false positives. On the held-out ckay16 test set the system achieves 96.0% accuracy, 95.7% TPR, 3.8% FPR, macro-F1 of 96.0%, and ROC-AUC of 0.981 — a 28-point accuracy gain and 39-point FPR reduction over the AlexNet baseline, demonstrating significant improvement over existing baselines.},
keywords = {Road accident detection, EfficientNet, CBAM, Attention mechanism, Focal loss, Test-time augmentation, Transfer learning, CCTV surveillance},
month = {April},
doi = {https://doi.org/10.64388/IREV9I10-1717084}
}