Home / Current Issue / Paper 1715144
Vision Mate: An AI Based Mobile Assistive System for Visually Impaired Users
Subject area: Arts, Social Sciences and Humanities · Area of research: Solapur
Abstract
VisionMate is a mobile accessibility application that utilizes artificial intelligence (AI) and helps visually impaired and blind people to carry out daily tasks independently. The device proposed integrates computer vision, speech recognition, and GPS to utilize the camera, microphone, and location services on a user's smartphone to assist users with a variety of tasks. VisionMate also provides a multitude of functions for people with visual impairments, including Optical Character Recognition (OCR) for reading text, Object Detection for finding people, currency detection for identifying denomination, Barcode Scanning to read product/food labels, and Voice Navigation for providing direction to users via voice command. Uses of cloud services (Google Vision API & Speech-to-Text API) to store incoming images and voice commands allow for the analysis of received images and voice commands. The image results are returned to the user via real-time sound using Text-to-Speech (TTS) technology. VisionMate will be developed using the React Native framework, and MongoDB will serve as the backend database for user management and authentication. Ultimately the goal of the proposed system is to improve access to and independence for visually impaired users by allowing users to better understand the environment using intelligent audio guidance.
Keywords
Artificial Intelligence, Assistive Technology, Computer Vision, Mobile Accessibility, Optical Character Recognition (OCR)
References
[1] D. Dakopoulos and N. G. Bourbakis, “Wearable obstacle avoidance electronic travel aids for blind: A survey,” IEEE Transactions on Systems, Man, and Cybernetics Part C, vol. 40, no. 1, pp. 25–35, 2010.
[2] M. Manduchi and J. Coughlan, “(Computer) Vision without sight,” Communications of the ACM, vol. 55, no. 1, pp. 96–104, 2012.
[3] J. M. Coughlan and H. Shen, “A smartphone-based system for assisting visually impaired users to navigate complex indoor environments,” IEEE International Conference on Systems, Man and Cybernetics, 2013.
[4] A. Ahmetovic, C. Gleason, K. Kitani, H. Takagi, and C. Asakawa, “NavCog: A navigation system for visually impaired users,” ACM SIGACCESS Conference on Computers and Accessibility, 2016.
[5] S. Mascetti, A. Ahmetovic, C. Bernareggi, and A. Gerino, “TypeInBraille: A braille-based typing interface for touchscreen devices,” ACM Transactions on Accessible Computing, 2016.
[6] J. Bigham et al., “VizWiz: Nearly real-time answers to visual questions,” Proceedings of the ACM Symposium on User Interface Software and Technology, 2010.
[7] A. Kane, J. Wobbrock, and R. Ladner, “Usable gestures for blind people: Understanding preference and performance,” ACM SIGCHI Conference on Human Factors in Computing Systems, 2011.
[8] S. Shoval, J. Borenstein, and Y. Koren, “Mobile robot obstacle avoidance in a computerized travel aid for the blind,” IEEE Transactions on Robotics and Automation, vol. 14, no. 3, pp. 485–492, 1998.
[9] J. M. Coughlan and A. Yuille, “The Manhattan world assumption: Regularities in scene statistics,” Neural Computation, vol. 15, no. 4, pp. 833–854, 2003.
[10] M. A. Goodrich and A. C. Schultz, “Human–robot interaction: A survey,” Foundations and Trends in Human–Computer Interaction, vol. 1, no. 3, pp. 203–275, 2007.
[11] R. Smith, “An overview of the Tesseract OCR engine,” Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), 2007.
[12] A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” Advances in Neural Information Processing Systems, 2012.
[13] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You Only Look Once: Unified real-time object detection,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
[14] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017.
[15] A. Howard et al., “MobileNets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
How to cite this paper
@article{1715144,
author = {Sarthak Maruti More, Vijayraj Dilip Chavan, Abhishek Dayanand Kundkar, Sohel Shabbir Nadaf, Megha Mukundrao Chaitanya},
title = {Vision Mate: An AI Based Mobile Assistive System for Visually Impaired Users},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {9},
pages = {1252-1258},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1715144.pdf},
abstract = {VisionMate is a mobile accessibility application that utilizes artificial intelligence (AI) and helps visually impaired and blind people to carry out daily tasks independently. The device proposed integrates computer vision, speech recognition, and GPS to utilize the camera, microphone, and location services on a user's smartphone to assist users with a variety of tasks. VisionMate also provides a multitude of functions for people with visual impairments, including Optical Character Recognition (OCR) for reading text, Object Detection for finding people, currency detection for identifying denomination, Barcode Scanning to read product/food labels, and Voice Navigation for providing direction to users via voice command. Uses of cloud services (Google Vision API & Speech-to-Text API) to store incoming images and voice commands allow for the analysis of received images and voice commands. The image results are returned to the user via real-time sound using Text-to-Speech (TTS) technology. VisionMate will be developed using the React Native framework, and MongoDB will serve as the backend database for user management and authentication. Ultimately the goal of the proposed system is to improve access to and independence for visually impaired users by allowing users to better understand the environment using intelligent audio guidance.},
keywords = {Artificial Intelligence, Assistive Technology, Computer Vision, Mobile Accessibility, Optical Character Recognition (OCR)},
month = {March},
doi = {https://doi.org/10.64388/IREV9I9-1715144}
}