International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1714023

1714023 Vol 9 · Issue 8 Download Paper

Praxis: A Hybrid Mobile AI Application with Offline LLM Inference and Cloud API Integration

Sri Jaya Lingeswaran A Keithick R Niranjan R Adrash S

Subject area: Science,Engineering and Technology  ·  Area of research: Artificial Intelligence

DOI: https://doi.org/10.64388/IREV9I8-1714023

Abstract

Large Language Models (LLMs) have revolutionized natural language processing and artificial intelligence applications. However, deploying these models on mobile devices presents significant challenges due to limited computational resources, memory constraints, and energy consumption. This paper introduces Praxis, a novel hybrid mobile application architecture that seamlessly integrates on-device LLM inference with cloud-based API services. Praxis enables users to run small language models locally on Android devices for enhanced privacy and offline capability while providing the flexibility to switch to more powerful cloud-based models when connectivity and computational demands permit. Our implementation leverages React Native, Expo, and integration with Hugging Face model repositories, combined with secure API key management for services like OpenAI, Anthropic Claude, and Google Gemini. Experimental results demonstrate that Praxis achieves a balance between performance, privacy, energy efficiency, and user experience, making advanced AI capabilities accessible on resourceconstrained mobile devices.

Keywords

Large Language Models, Mobile Computing, On-Device Inference, Edge Computing, Hybrid Architecture, Privacy-Preserving AI, React Native

References

[1] Author et al., “Mobile Edge Intelligence for Large Language Models: A Contemporary Survey,” IEEE Wireless Communications, 2025. DOI: 10.1109/xxx.2025.10835069

[2] Author et al., “m-LLM: A Multi-Dimensional Optimization Framework for LLM Inference on Mobile Devices,” IEEE Transactions on Mobile Computing, 2024. DOI: 10.1109/TMC.2024.11075620

[3] Author et al., “EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,” IEEE Transactions on Computers, 2025. DOI: 10.1109/TC.2025.10812936

[4] Author et al., “Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM,” Proc. IEEE International Conference on Computer Design (ICCD), 2024. DOI: 10.1109/ICCD.2024.10764574

[5] Author et al., “Large Language Models (LLMs) Inference Offloading and Resource Allocation in Cloud-Edge Networks: An Active Inference Approach,” Proc. IEEE Global Communications Conference (GLOBECOM), 2023. DOI: 10.1109/GLOBECOM.2023.10333824

[6] Author et al., “FACIL: Flexible DRAM Address Mapping for SoCPIM Cooperative On-device LLM Inference,” Proc. IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2025. DOI: 10.1109/HPCA.2025.10946824

[7] Author et al., “EdgeShard: Efficient LLM Inference via Collaborative Edge Computing,” IEEE Transactions on Mobile Computing, 2025. DOI: 10.1109/TMC.2025.10818760

[8] Author et al., “Generative Inference of Large Language Models in Edge Computing: An Energy Efficient Approach,” Proc. IEEE International Conference on Web Services (ICWS), 2024. DOI: 10.1109/ICWS.2024.10592339

[9] Author et al., “Hybrid AI Architecture using Edge-Cloud Computing for Secure V2X Communication,” Proc. IEEE Vehicular Technology Conference (VTC), 2024. DOI: 10.1109/VTC.2024.10859430

[10] Author et al., “A hierarchical edge cloud architecture for mobile computing,” Proc. IEEE International Conference on Computer Communications (INFOCOM), 2016. DOI: 10.1109/INFOCOM.2016.7524340

[11] Author et al., “Edge-Cloud Computing and Artificial Intelligence in Internet of Medical Things: Architecture, Technology and Application,” IEEE Internet of Things Journal, vol. 7, no. 8, pp. 6689-6706, 2020. DOI: 10.1109/JIOT.2020.2999359

[12] Author et al., “DeePar: A Hybrid Device-Edge-Cloud Execution Framework for Mobile Deep Learning Applications,” Proc. IEEE International Conference on Web Services (ICWS), 2019. DOI: 10.1109/ICWS.2019.8845240

[13] Author et al., “Cloud-Edge Orchestration for the Internet of Things: Architecture and AI-Powered Data Processing,” IEEE Internet of ThingsJournal,vol.7,no. 10, pp.9847-9859,2020.DOI: 10.1109/JIOT.2020.3009809

How to cite this paper

Sri Jaya Lingeswaran A, Keithick R, Niranjan R, Adrash S "Praxis: A Hybrid Mobile AI Application with Offline LLM Inference and Cloud API Integration" Iconic Research And Engineering Journals Volume 9 Issue 8 2026 Page 2555-2559 https://doi.org/10.64388/IREV9I8-1714023
Sri Jaya Lingeswaran A, Keithick R, Niranjan R, Adrash S "Praxis: A Hybrid Mobile AI Application with Offline LLM Inference and Cloud API Integration" Iconic Research And Engineering Journals, vol. 9, no. 8, Feb. 2026, doi: https://doi.org/10.64388/IREV9I8-1714023
Sri Jaya Lingeswaran A, Keithick R, Niranjan R, Adrash S (2026). Praxis: A Hybrid Mobile AI Application with Offline LLM Inference and Cloud API Integration. Iconic Research And Engineering Journals, 9(8). doi: https://doi.org/10.64388/IREV9I8-1714023
Sri Jaya Lingeswaran A, Keithick R, Niranjan R, Adrash S "Praxis: A Hybrid Mobile AI Application with Offline LLM Inference and Cloud API Integration" Iconic Research And Engineering Journals, vol. 9, no. 8, Feb. 2026. Crossref, https://doi.org/10.64388/IREV9I8-1714023
@article{1714023,
      author = {Sri Jaya Lingeswaran A, Keithick R, Niranjan R, Adrash S},
      title = {Praxis: A Hybrid Mobile AI Application with Offline LLM Inference and Cloud API Integration},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {8},
      pages = {2555-2559},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1714023.pdf},
      abstract = {Large Language Models (LLMs) have revolutionized natural language processing and artificial intelligence applications. However, deploying these models on mobile devices presents significant challenges due to limited computational resources, memory constraints, and energy consumption. This paper introduces Praxis, a novel hybrid mobile application architecture that seamlessly integrates on-device LLM inference with cloud-based API services. Praxis enables users to run small language models locally on Android devices for enhanced privacy and offline capability while providing the flexibility to switch to more powerful cloud-based models when connectivity and computational demands permit. Our implementation leverages React Native, Expo, and integration with Hugging Face model repositories, combined with secure API key management for services like OpenAI, Anthropic Claude, and Google Gemini. Experimental results demonstrate that Praxis achieves a balance between performance, privacy, energy efficiency, and user experience, making advanced AI capabilities accessible on resourceconstrained mobile devices.},
      keywords = {Large Language Models, Mobile Computing, On-Device Inference, Edge Computing, Hybrid Architecture, Privacy-Preserving AI, React Native},
      month = {February},
      doi = {https://doi.org/10.64388/IREV9I8-1714023}
  }