Home / Current Issue / Paper 1714023
Praxis: A Hybrid Mobile AI Application with Offline LLM Inference and Cloud API Integration
Subject area: Science,Engineering and Technology · Area of research: Artificial Intelligence
DOI: https://doi.org/10.64388/IREV9I8-1714023
Abstract
Large Language Models (LLMs) have revolutionized natural language processing and artificial intelligence applications. However, deploying these models on mobile devices presents significant challenges due to limited computational resources, memory constraints, and energy consumption. This paper introduces Praxis, a novel hybrid mobile application architecture that seamlessly integrates on-device LLM inference with cloud-based API services. Praxis enables users to run small language models locally on Android devices for enhanced privacy and offline capability while providing the flexibility to switch to more powerful cloud-based models when connectivity and computational demands permit. Our implementation leverages React Native, Expo, and integration with Hugging Face model repositories, combined with secure API key management for services like OpenAI, Anthropic Claude, and Google Gemini. Experimental results demonstrate that Praxis achieves a balance between performance, privacy, energy efficiency, and user experience, making advanced AI capabilities accessible on resourceconstrained mobile devices.
Keywords
Large Language Models, Mobile Computing, On-Device Inference, Edge Computing, Hybrid Architecture, Privacy-Preserving AI, React Native
How to cite this paper
@article{1714023,
author = {Sri Jaya Lingeswaran A, Keithick R, Niranjan R, Adrash S},
title = {Praxis: A Hybrid Mobile AI Application with Offline LLM Inference and Cloud API Integration},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {8},
pages = {2555-2559},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1714023.pdf},
abstract = {Large Language Models (LLMs) have revolutionized natural language processing and artificial intelligence applications. However, deploying these models on mobile devices presents significant challenges due to limited computational resources, memory constraints, and energy consumption. This paper introduces Praxis, a novel hybrid mobile application architecture that seamlessly integrates on-device LLM inference with cloud-based API services. Praxis enables users to run small language models locally on Android devices for enhanced privacy and offline capability while providing the flexibility to switch to more powerful cloud-based models when connectivity and computational demands permit. Our implementation leverages React Native, Expo, and integration with Hugging Face model repositories, combined with secure API key management for services like OpenAI, Anthropic Claude, and Google Gemini. Experimental results demonstrate that Praxis achieves a balance between performance, privacy, energy efficiency, and user experience, making advanced AI capabilities accessible on resourceconstrained mobile devices.},
keywords = {Large Language Models, Mobile Computing, On-Device Inference, Edge Computing, Hybrid Architecture, Privacy-Preserving AI, React Native},
month = {February},
doi = {https://doi.org/10.64388/IREV9I8-1714023}
}