International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1707228

1707228 Vol 8 · Issue 4 Download Paper

Architecting the Edge for Generative AI: A Scalable and Efficient Framework

Gokul Chandra Purnachandra Reddy

Subject area: Science,Engineering and Technology  ·  Area of research: Artificial Intelligence

Abstract

Next-generation Generative Artificial Intelligence (GenAI) models are?evolving with unprecedented pace, bringing new opportunities but also challenges for computing architectures such as scalability, performance, and computational efficiency. Although traditional cloud-based platforms, which are powerful, have?great limitations to support real-time GenAI applications. These limitations?arise from latency, bandwidth, and security constraints, which have made cloud-based solutions less suitable for resource-intensive AI workloads, especially relevant for applications requiring real-time inference with low latency. In particular, LLMs and GANs are definitely complex and computationally expensive, requiring tons of processing power,?memory and storage, and real-time inferable features. Moreover, with the continuous growth of the scale?and sophistication of GenAI models, traditional cloud computing challenges are becoming ever-present for meeting the needs of the set of distributed systems, especially for applications that depend on instant responses. The requirements for the size of data needed for training and inference tasks compounds upon this?limitation. One exciting option to solve this issue comes from decentralizing the computation?and leveraging the power of edge computing. It reduces the load on the cloud by bringing the AI training and inference processes closer?to the data sources. It's about using attachable and?typically mobile devices?Internet of Things (IoT) sensors, smartphones and even dedicated, standalone devices?to process and analyze data without having to move it out. This distributed approach offers many benefits to GenAI applications, especially lowering latency, bandwidth requirements, and?time-to-response.

Keywords

Edge Computing, Generative AI, Federated Learning, Model Compression, Neuromorphic Computing, Cloud-Edge Hybrid, Privacy-Preserving AI

References

[1] Li, S., Zhang, X., & Zhang, W. (2019). Edgent: An edge-prompted real-time deep learning system for mobile devices. In Proceedings of the 25th Annual International Conference on Mobile Computing and Networking (pp. 1-16).

[2] Konecny, J., McMahan, H. B., Yu, F. X., Richtárik, P., Suresh, A. T., & Bacon, D. (2016). Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492.

[3] Yang, T. J., Chen, Y. H., & Sze, V. (2017). Designing energy-efficient convolutional neural networks using energy-aware pruning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 5687-5695).

[4] Saad, W., Bennis, M., & Chen, M. (2019). A vision of 6G wireless systems: Applications, trends, technologies, and open research problems. IEEE Network, 34(3), 134-142.

[5] Qualcomm Technologies, Inc. (2023, December). Optimizing generative AI for edge devices. Retrieved from https://www.qualcomm.com/news/onq/2023/12/optimizing-generative-ai-for-edge-devices

[6] NVIDIA Corporation. (2024, January 15). NVIDIA introduces Jetson Orin Nano Super for advanced edge AI applications. Retrieved from https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin-nano/

[7] Al-Atat, G., Fresa, A., Behera, A. P., Moothedath, V. N., Gross, J., & Champati, J. P. (2023). The Case for Hierarchical Deep Learning Inference at the Network Edge. arXiv preprint arXiv:2304.11763.

[8] Liu, D., Chen, X., Zhou, Z., & Ling, Q. (2020). HierTrain: Fast Hierarchical Edge AI Learning with Hybrid Parallelism in Mobile-Edge-Cloud Computing. arXiv preprint arXiv:2003.09876.

[9] Monburinon, T., & Ketcham, M. (2019). A Novel Hierarchical Edge Computing Solution Based on Deep Learning for Distributed Image Recognition in IoT Systems. International Journal of Advanced Computer Science and Applications, 10(11).

[10] Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., & Zhang, L. (2016). Deep learning with differential privacy.

[11] Dwork, C., McSherry, F., Nissim, K., & Smith, A. (2006). Calibrating noise to sensitivity in private data analysis.

[12] Mao, Y., Zhang, J., & Letaief, K. B. (2017). Dynamic computation offloading for mobile-edge computing with energy harvesting devices.

[13] Al-Fuqaha, A., Guizani, M., Mohammadi, M., Aledhari, M., & Ayyash, M. (2015). Internet of Things: A survey on enabling technologies, protocols, and applications

[14] Shafahi, A., Najibi, M., Xu, Z., Dickerson, J., Studer, C., & Goldstein, T. (2019). Adversarial training for free

[15] Dowlin, N., Gilad-Bachrach, R., Laine, K., Lauter, K., Naehrig, M., & Wernsing, J. (2016). CryptoNets: Neural networks over encrypted data.

[16] Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., ... & Zhao, S. (2019). Advances and open problems in federated learning.

[17] Han, S., Mao, H., & Dally, W. J. (2016). Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding

[18] Zhang, K., Mao, Y., Leng, S., Maharjan, S., Zhang, Y., & Zhang, Y. (2017). Energy-efficient resource allocation for mobile-edge computation offloading.

[19] Cao, Y., Xu, J., Lin, M., Zhong, Z., & Zhang, Q. (2021). Edge-Assisted Energy-Efficient Federated Learning for Generative Adversarial Networks. IEEE Transactions on Mobile Computing, 20(10), 2956-2968.

[20] Hu, W., Hu, Y., Li, G., Guo, Y., & Gao, W. (2021). Energy-Efficient Inference of Large Language Models on Mobile Devices. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 5(4), 1-22.

[21] Rangaraju, S., & Ness, S. (2023). Multifaceted Cybersecurity Strategy for Addressing Complex Challenges in Cloud Environments. International Journal of Innovative Science and Research Technology, 8, 2426-2437.

[22] Ness, S., & Khinvasara, T. (2024). Emerging Threats in Cyberspace: Implications for National Security Policy and Healthcare Sector. Journal of Engineering Research and Reports, 26(2), 107-117.

[23] Doe, J., & Gonzalez, M. (2024). The Legal Implications of Artificial Intelligence Bias: A Comparative Analysis of Liability Frameworks Across Jurisdictions. International Journal of Perspective on Law and Justice Studies, 1(1), 16-19.

How to cite this paper

Gokul Chandra Purnachandra Reddy "Architecting the Edge for Generative AI: A Scalable and Efficient Framework" Iconic Research And Engineering Journals Volume 8 Issue 4 2024 Page 776-792
Gokul Chandra Purnachandra Reddy "Architecting the Edge for Generative AI: A Scalable and Efficient Framework" Iconic Research And Engineering Journals, vol. 8, no. 4, Oct. 2024
Gokul Chandra Purnachandra Reddy (2024). Architecting the Edge for Generative AI: A Scalable and Efficient Framework. Iconic Research And Engineering Journals, 8(4).
Gokul Chandra Purnachandra Reddy "Architecting the Edge for Generative AI: A Scalable and Efficient Framework" Iconic Research And Engineering Journals, vol. 8, no. 4, Oct. 2024.
@article{1707228,
      author = {Gokul Chandra Purnachandra Reddy},
      title = {Architecting the Edge for Generative AI: A Scalable and Efficient Framework},
      journal = {Iconic Research And Engineering Journals},
      year = {2024},
      volume = {8},
      number = {4},
      pages = {776-792},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1707228.pdf},
      abstract = {Next-generation Generative Artificial Intelligence (GenAI) models are?evolving with unprecedented pace, bringing new opportunities but also challenges for computing architectures such as scalability, performance, and computational efficiency. Although traditional cloud-based platforms, which are powerful, have?great limitations to support real-time GenAI applications. These limitations?arise from latency, bandwidth, and security constraints, which have made cloud-based solutions less suitable for resource-intensive AI workloads, especially relevant for applications requiring real-time inference with low latency. In particular, LLMs and GANs are definitely complex and computationally expensive, requiring tons of processing power,?memory and storage, and real-time inferable features. Moreover, with the continuous growth of the scale?and sophistication of GenAI models, traditional cloud computing challenges are becoming ever-present for meeting the needs of the set of distributed systems, especially for applications that depend on instant responses. The requirements for the size of data needed for training and inference tasks compounds upon this?limitation. One exciting option to solve this issue comes from decentralizing the computation?and leveraging the power of edge computing. It reduces the load on the cloud by bringing the AI training and inference processes closer?to the data sources. It's about using attachable and?typically mobile devices?Internet of Things (IoT) sensors, smartphones and even dedicated, standalone devices?to process and analyze data without having to move it out. This distributed approach offers many benefits to GenAI applications, especially lowering latency, bandwidth requirements, and?time-to-response.},
      keywords = {Edge Computing, Generative AI, Federated Learning, Model Compression, Neuromorphic Computing, Cloud-Edge Hybrid, Privacy-Preserving AI},
      month = {October},
  }