International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1711192

1711192 Vol 6 · Issue 10 Download Paper

Energy-Efficient Ai Model Design for Edge Devices Using Neural Network Pruning and Optimization Techniques

Rishabh Agrawal Himanshu Kumar

Subject area: Science,Engineering and Technology  ·  Area of research: Artificial Intelligence and Edge Computing

DOI: https://doi.org/10.64388/IREV6I10-1711192

Abstract

AI is progressively being implemented at the edge computing platforms of smartphones, wearables, industrial sensors, and autonomous systems. Although such deployments can support real-time processing, preserve privacy, and decrease reliance on the network, they tend to be restricted by limited computational power, limited memory, and strict power requirements. Without much optimization, the traditional deep neural networks with their large number of parameters and high computational needs are ill-suited to such environments. The present article discusses the neural network pruning and complementary optimization methods as possible solutions to these issues by suggesting the energy-efficient design of AI models. Pruning is used to remove redundant parameters to reduce model size and operations, and quantization is used to encode high-precision weights into low-bit representations to reduce memory and energy usage. Efficiency is additionally improved with knowledge distillation and lightweight architectures with no performance costs, and compiler-level optimizations are applied to guarantee that compressed models can produce real-world runtime gains on a variety of hardware platforms. The discussion combines theoretical knowledge and practical processes, such as step-by-step design processes and example codes, and latency, memory footprint, and energy consumption measuring guidelines on actual models. Issues like accuracy loss, heterogeneity of hardware, and use of standardized benchmarks are critically discussed, and future research directions, including ultra-low-bit networks, hardware-aware neural architecture search, and energy-centric training objectives, are discussed. Through a combination of cutting-edge approaches and deployment-focused ideas, this piece of work highlights that AI minimal energy usage is not only a technical one but an important action towards sustainable, scaled, and accessible edge computing.

Keywords

Energy-Efficient AI; Edge Computing; Neural Network Pruning; Model Optimization; Quantization; Knowledge Distillation; Lightweight Architectures; Compiler Optimizations

References

[1] Han, S., Mao, H., & Dally, W. J. (2015). Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. arXiv preprint. Demonstrates a full pipeline combining pruning and quantization for substantial modelsize reduction and energy gains. arXiv Link

[2] De Leon, J. D., & Atienza, R. (2022). Depth Pruning with Auxiliary Networks for TinyML. arXiv preprint. Achieved up to 93% parameter reduction with minimal accuracy loss on TinyML tasks using depth pruning. Papers with CodearXiv

[3] Blakeney, C., Li, X., Yan, Y., & Zong, Z. (2020). Parallel Blockwise Knowledge Distillation for Deep Neural Network Compression. arXiv preprint. Shows 3× speedup and ~20–30% energy savings using blockwise distillation. arXiv Link

[4] Bharti, K., Cervera-Lierta, A., Kyaw, T. H., Haug, T., Alperin-Lea, S., Anand, A., . . . Aspuru-Guzik, A. (2022). Noisy intermediate-scale quantum algorithms. Reviews of Modern Physics, 94(1). https://doi.org/10.1103/revmodphys.94.015004

[5] Davies, M., Wild, A., Orchard, G., Sandamirskaya, Y., Guerra, G. a. F., Joshi, P., . . .

[6] Risbud, S. R. (2021). Advancing Neuromorphic Computing with LOIHI: A Survey of Results and Outlook. Proceedings of the IEEE, 109(5), 911–934. https://doi.org/10.1109/jproc.2021.3067593

[7] Mahdavinejad, M. S., Rezvan, M., Barekatain, M., Adibi, P., Barnaghi, P., & Sheth, A. P. (2017). Machine learning for internet of things data analysis: a survey. Digital Communications and Networks, 4(3), 161–175. https://doi.org/10.1016/j.dcan.2017.10.002

[8] Sze, V., Chen, Y., Yang, T., & Emer, J. S. (2017). Efficient Processing of deep Neural Networks: A tutorial and survey. Proceedings of the IEEE, 105(12), 2295–2329. https://doi.org/10.1109/jproc.2017.2761740

[9] Wang, X., Han, Y., Leung, V. C. M., Niyato, D., Yan, X., & Chen, X. (2020). Convergence of Edge Computing and Deep Learning: A Comprehensive survey. IEEE Communications Surveys & Tutorials, 22(2), 869–904. https://doi.org/10.1109/comst.2020.2970550

[10] Xin, Y., Kong, L., Liu, Z., Chen, Y., Li, Y., Zhu, H., . . . Wang, C. (2018). Machine learning and deep learning methods for cybersecurity. IEEE Access, 6, 35365–35381. https://doi.org/10.1109/access.2018.2836950

[11] Zhang, C., Patras, P., & Haddadi, H. (2019). Deep learning in mobile and wireless Networking: a survey. IEEE Communications Surveys & Tutorials, 21(3), 2224–2287. https://doi.org/10.1109/comst.2019.2904897

[12] Zhou, Z., Chen, X., Li, E., Zeng, L., Luo, K., & Zhang, J. (2019). Edge Intelligence: Paving the last mile of artificial intelligence with edge computing. Proceedings of the IEEE, 107(8), 1738–1762. https://doi.org/10.1109/jproc.2019.2918951

[13] Brunetti, A., Buongiorno, D., Trotta, G. F., & Bevilacqua, V. (2018). Computer vision and deep learning techniques for pedestrian detection and tracking: A survey. Neurocomputing, 300, 17–33. https://doi.org/10.1016/j.neucom.2018.01.092

[14] Deng, S., Zhao, H., Fang, W., Yin, J., Dustdar, S., & Zomaya, A. Y. (2020). Edge Intelligence: The confluence of edge computing and artificial intelligence. IEEE Internet of Things Journal, 7(8), 7457–7469. https://doi.org/10.1109/jiot.2020.2984887

[15] Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., & Keutzer, K. (2022). A survey of Quantization Methods for Efficient Neural network Inference. In Chapman and Hall/CRC eBooks (pp. 291–326). https://doi.org/10.1201/9781003162810-13

[16] Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., . . . Zhao, S. (2021). Advances and open problems in federated learning. https://doi.org/10.1561/9781680837896

[17] Letaief, K. B., Shi, Y., Lu, J., & Lu, J. (2021). Edge Artificial Intelligence for 6G: vision, enabling technologies, and applications. IEEE Journal on Selected Areas in Communications, 40(1), 5–36. https://doi.org/10.1109/jsac.2021.3126076

[18] Xu, J., Glicksberg, B. S., Su, C., Walker, P., Bian, J., & Wang, F. (2020). Federated Learning for Healthcare Informatics. Journal of Healthcare Informatics Research, 5(1), 1–19. https://doi.org/10.1007/s41666-020-00082-4

[19] Courbariaux, M., Bengio, Y., & David, J. P. (2015). BinaryConnect: Training deep neural networks with binary weights during propagations. In Advances in Neural Information Processing Systems (NeurIPS), 28.

[20] Rastegari, M., Ordonez, V., Redmon, J., & Farhadi, A. (2016). XNOR-Net: ImageNet classification using binary convolutional neural networks. In European Conference on Computer Vision (ECCV) (pp. 525–542). Springer.

[21] Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., ... & Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861.

[22] Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L. C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4510–4520).

[23] Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., ... & Adam, H. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2704–2713).

[24] Yang, T. J., Chen, Y. H., & Sze, V. (2020). Designing energy-efficient convolutional neural networks using energy-aware pruning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5687–5695).

[25] Qiu, H., Wang, J., Chen, X., & Shen, Y. (2022). Recent advances in neural network compression and acceleration for edge AI. ACM Computing Surveys (CSUR), 54(9), 1–36.

How to cite this paper

Rishabh Agrawal, Himanshu Kumar "Energy-Efficient Ai Model Design for Edge Devices Using Neural Network Pruning and Optimization Techniques" Iconic Research And Engineering Journals Volume 6 Issue 10 2023 Page 1141-1155 https://doi.org/10.64388/IREV6I10-1711192
Rishabh Agrawal, Himanshu Kumar "Energy-Efficient Ai Model Design for Edge Devices Using Neural Network Pruning and Optimization Techniques" Iconic Research And Engineering Journals, vol. 6, no. 10, Apr. 2023, doi: https://doi.org/10.64388/IREV6I10-1711192
Rishabh Agrawal, Himanshu Kumar (2023). Energy-Efficient Ai Model Design for Edge Devices Using Neural Network Pruning and Optimization Techniques. Iconic Research And Engineering Journals, 6(10). doi: https://doi.org/10.64388/IREV6I10-1711192
Rishabh Agrawal, Himanshu Kumar "Energy-Efficient Ai Model Design for Edge Devices Using Neural Network Pruning and Optimization Techniques" Iconic Research And Engineering Journals, vol. 6, no. 10, Apr. 2023. Crossref, https://doi.org/10.64388/IREV6I10-1711192
@article{1711192,
      author = {Rishabh Agrawal, Himanshu Kumar},
      title = {Energy-Efficient Ai Model Design for Edge Devices Using Neural Network Pruning and Optimization Techniques},
      journal = {Iconic Research And Engineering Journals},
      year = {2023},
      volume = {6},
      number = {10},
      pages = {1141-1155},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1711192.pdf},
      abstract = {AI is progressively being implemented at the edge computing platforms of smartphones, wearables, industrial sensors, and autonomous systems. Although such deployments can support real-time processing, preserve privacy, and decrease reliance on the network, they tend to be restricted by limited computational power, limited memory, and strict power requirements. Without much optimization, the traditional deep neural networks with their large number of parameters and high computational needs are ill-suited to such environments. The present article discusses the neural network pruning and complementary optimization methods as possible solutions to these issues by suggesting the energy-efficient design of AI models. Pruning is used to remove redundant parameters to reduce model size and operations, and quantization is used to encode high-precision weights into low-bit representations to reduce memory and energy usage. Efficiency is additionally improved with knowledge distillation and lightweight architectures with no performance costs, and compiler-level optimizations are applied to guarantee that compressed models can produce real-world runtime gains on a variety of hardware platforms. The discussion combines theoretical knowledge and practical processes, such as step-by-step design processes and example codes, and latency, memory footprint, and energy consumption measuring guidelines on actual models. Issues like accuracy loss, heterogeneity of hardware, and use of standardized benchmarks are critically discussed, and future research directions, including ultra-low-bit networks, hardware-aware neural architecture search, and energy-centric training objectives, are discussed. Through a combination of cutting-edge approaches and deployment-focused ideas, this piece of work highlights that AI minimal energy usage is not only a technical one but an important action towards sustainable, scaled, and accessible edge computing.},
      keywords = {Energy-Efficient AI; Edge Computing; Neural Network Pruning; Model Optimization; Quantization; Knowledge Distillation; Lightweight Architectures; Compiler Optimizations},
      month = {April},
      doi = {https://doi.org/10.64388/IREV6I10-1711192}
  }