Home / Current Issue / Paper 1712742
Real-Time AI-Driven Exam Cheating Detection Using YOLO and Pose Estimation: A Multi-Modal Deep Learning Approach
Subject area: Science,Engineering and Technology · Area of research: Artificial Intelligence
Abstract
The preservation of academic integrity in examination environments is a cornerstone of credible certification. Traditional invigilation methods, reliant on human monitors, are inherently limited by factors such as fatigue, subjective bias, and an inability to monitor large-scale settings simultaneously. This paper proposes a novel, end-to-end, real-time AI-driven system for the automated detection of exam cheating. Our solution employs a multi-modal framework that combines state-of-the-art object detection with sophisticated human pose estimation. We implement and fine-tune a YOLOv8 (You Only Look Once) model for the real-time identification of prohibited objects, including mobile phones, micro-earpieces, and written notes. Concurrently, an OpenPose-based pose estimation pipeline analyzes candidates' body language to flag suspicious postures, such as excessive head rotation for gaze estimation, abnormal body orientation, and furtive hand movements. A central decision module fuses these dual streams of visual evidence using a rule-based heuristic to generate low-latency, high-confidence alerts for human invigilators. To validate our system, we curated a comprehensive custom dataset comprising 50 hours of annotated exam footage. Experimental results reveal that our fused model achieves a precision of 0.92, a recall of 0.88, and F1-score of 0.90, significantly outperforming unimodal baselines (YOLO-only and Pose-only). The system operates at an average latency of 35 ms per frame, fulfilling the stringent requirements for real-time video surveillance. This work establishes a robust, scalable, and efficient paradigm for proactive exam integrity enforcement, mitigating the limitations of human-based monitoring.
Keywords
AI-driven, Object Detection, YOLOv8, Human Pose Estimation, Multi-Modal Fusion, Real-Time Surveillance, Deep Learning
References
[1] Bochkovskiy, A., Wang, C. Y., & Liao, H. Y. M. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv preprint arXiv:2004.10934.
[2] Cao, Z., Hidalgo, G., Simon, T., Wei, S. E., & Sheikh, Y. (2021). OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(1), 172-186.
[3] Cheng, K., Wang, J., & Li, Q. (2021). A Survey on AI-based Online Exam Proctoring. Journal of Artificial Intelligence Research, 71, 1-35.
[4] Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 580-587).
[5] Jocher, G., Chaurasia, A. and Qiu, J. (2023). YOLO by Ultralytics. https://github.com/ultralytics/ultralytics
[6] Kumar, A., & Lee, S. (2023). Multi-modal Learning for Automated Proctoring: A Review. ACM Computing Surveys. (Placeholder)
[7] Newell, A., Yang, K., & Deng, J. (2016). Stacked hourglass networks for human pose estimation. European conference on computer vision (pp. 483-499).
[8] Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 779-788).
[9] Redmon, J., & Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv preprint arXiv:1804.02767.
[10] Sarsa, S., et al. (2022). Lightweight Object Detection for Mobile Exam Proctoring. Proceedings of the IEEE International Conference on EdTech.
[11] Tiong, L. C. O., & Lee, H. J. (2021). The cognitive load of human proctors in large-scale online examinations. Computers & Education, 164, 104121.
[12] Toshev, A., & Szegedy, C. (2014). Deeppose: Human pose estimation via deep neural networks. Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1653-1660).
[13] Ullah, A., et al. (2021). A Pose-based Anomaly Detection System for Exam Monitoring. Journal of Visual Communication and Image Representation, 79, 103-112. (Placeholder)
[14] Zhang, L., & Wang, H. (2020). The Challenges and Opportunities of Intelligent Surveillance. IEEE International Conference on Multimedia and Expo (ICME). (Placeholder)
[15] Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., & Ren, D. (2020). Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression. Proceedings of the AAAI Conference on Artificial Intelligence, 34(07), 12993-13000.
How to cite this paper
@article{1712742,
author = {Abdulrahman Abdulkarim, Muhammed Kuliya, Aisha Bappa Adam},
title = {Real-Time AI-Driven Exam Cheating Detection Using YOLO and Pose Estimation: A Multi-Modal Deep Learning Approach},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {6},
pages = {847-856},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1712742.pdf},
abstract = {The preservation of academic integrity in examination environments is a cornerstone of credible certification. Traditional invigilation methods, reliant on human monitors, are inherently limited by factors such as fatigue, subjective bias, and an inability to monitor large-scale settings simultaneously. This paper proposes a novel, end-to-end, real-time AI-driven system for the automated detection of exam cheating. Our solution employs a multi-modal framework that combines state-of-the-art object detection with sophisticated human pose estimation. We implement and fine-tune a YOLOv8 (You Only Look Once) model for the real-time identification of prohibited objects, including mobile phones, micro-earpieces, and written notes. Concurrently, an OpenPose-based pose estimation pipeline analyzes candidates' body language to flag suspicious postures, such as excessive head rotation for gaze estimation, abnormal body orientation, and furtive hand movements. A central decision module fuses these dual streams of visual evidence using a rule-based heuristic to generate low-latency, high-confidence alerts for human invigilators. To validate our system, we curated a comprehensive custom dataset comprising 50 hours of annotated exam footage. Experimental results reveal that our fused model achieves a precision of 0.92, a recall of 0.88, and F1-score of 0.90, significantly outperforming unimodal baselines (YOLO-only and Pose-only). The system operates at an average latency of 35 ms per frame, fulfilling the stringent requirements for real-time video surveillance. This work establishes a robust, scalable, and efficient paradigm for proactive exam integrity enforcement, mitigating the limitations of human-based monitoring.},
keywords = {AI-driven, Object Detection, YOLOv8, Human Pose Estimation, Multi-Modal Fusion, Real-Time Surveillance, Deep Learning},
month = {December},
doi = {https://doi.org/10.64388/IREV9I6-1712742}
}