International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1707828

1707828 Vol 8 · Issue 10 Download Paper

A Survey Comparing Specialized Hardware and Evolution in CPU, GPU AND TPU for Neural Network

Francis Chigozie Odikwa Ndubuisi Onyemaobi Bethram Chibuzo

Subject area: Science,Engineering and Technology  ·  Area of research: Science

Abstract

This survey paper is focused on the genesis of GPU AND TPU from first generation and their architectures. The paper compares CPUs, GPUs, and TPUs, their hardware architectures, their similarities and differences were extensively discussed. The batch of data are immensely used these days but they require more time, computation and energy. Due to the greater demand and attractive options for architects to explore, companies are continuously working to reduce training and inference response time. Pertinent to the demands and cost factors different kinds of ASICs (application specific integrated circuits) are developed and research is increased in this area. Many models of CPUs, GPUs and TPUs have been developed to support these networks and to improve training and inference phase. The hardware of CPUs and GPUs can be sold to businesses while Google offers TPU processing for everyone from the cloud. When the data is away from the computational source, it increases the overall cost and to reduce this cost companies implements memory management and caching techniques close to ALUs.

Keywords

GPU, TPU, CPU, Deep Learning, ALU, framework.

References

[1] “TensorFlow: Using JIT compilation https://www.tensorflow.org/xla/jit,” 2018.

[2] “https://cloud.google.com/tpu/docs/system-architecture,” Google Cloud Documentation, 2018.

[3] Adolf, R. Rama, S. Reagen, B., and Brooks, D. “Fathom: Reference workloads for modern deep learning methods,” in Workload Characterization (IISWC), 2016 IEEE International Symposium on. IEEE, 2016, pp. 1–10.

[4] Amodei, D. Ananthanarayanan, S. Anubhai, R. Bai, J. Battenberg, E. Case, C. Casper, J. Catanzaro, B. Cheng, Q. and Chen G., “Deep speech 2: End-to-end speech recognition in english and mandarin,” in International Conference on Machine Learning, 2016, pp. 173–182.

[5] Bahrampour, S. Ramakrishnan, N. Schott, L. and Shah, M. “Comparative study of Caffe, Neon, Theano, and Torch for deep learning,” in ICLR, 2016.

[6] Banner, R. Hubara, I. Hoffer, E. and Soudry, D. “Scalable methods for 8-bit training of neural networks,” arXiv preprint arXiv:1805.11046, 2018.

[7] Bienia, C. Kumar, S. Singh, J. P. and Li, K. “The PARSEC benchmark suite: Characterization and architectural implications,” in Proceedings of the 17th international conference on Parallel architectures and compilation techniques. ACM, 2008, pp. 72–81.

[8] Blog, G. A. “Introducing GPipe, an open source library for efficiently training large-scale neural network models,” https://ai.googleblog.com/2019/03/introducing-gpipe-open-sourcelibrary.html, 2019.

[9] Case, L. “Volta Tensor Core GPU achieves new AI performance milestones,” Nvidia Developer Blog, 2018.

[10] Che, S. Boyer, M. Meng, J. Tarjan, D. Sheaffer, J. W. Lee, S.-H. and Skadron, K. “Rodinia: A benchmark suite for heterogeneous computing,” in Workload Characterization, 2009. IISWC 2009. IEEE International Symposium on. Ieee, 2009, pp. 44–54.

[11] Chen, T. Moreau, T. Jiang, Z. Shen, H. Yan, E. Wang, L. Hu, Y. Ceze, L. Guestrin, C. and Krishnamurthy, A. “TVM: End-to-end optimization stack for deep learning,” arXiv preprint arXiv:1802.04799, 2018.

[12] Chen, T. Chen, Y. Duranton, M. Guo, Q. Hashmi, A. Lipasti, M. Nere, A. Qiu, S. ebag, M. S and Temam, O. “BenchNN: On the broad potential application scope of hardware neural network accelerators,” in Workload Characterization (IISWC), 2012 IEEE International Symposium on.IEEE, 2012, pp. 36–45.

[13] Onyemaobi, C.B. and Ajah, I.A. (2017) “Comparative analysis of web development languages performances”, Int. J. Web Science, Vol. 3, No. 1, pp.16–31.

[14] Onyemaobi B.C. (2012). Design and Implementation of Aerospace Information System. Lokoja: Lambert Academic Publishing.

[15] O. B. Chibuzo and D. O. Isiaka, “Design and Implementation of Secure Browser for ComputerBased Tests,” Int. J. Innov. Sci. Res. Technol., vol. 5, no. 8, pp. 1347–1356, 2020, doi: 10.38124/ijisrt20aug526.

[16] Bethram Chibuzo, O., Philip Omoniyi, A.: Network and complex systems developing a signal booster for improved communication in remote areas (2019). https://doi.org/10.7176/NCS

How to cite this paper

Francis Chigozie, Odikwa Ndubuisi, Onyemaobi Bethram Chibuzo "A Survey Comparing Specialized Hardware and Evolution in CPU, GPU AND TPU for Neural Network" Iconic Research And Engineering Journals Volume 8 Issue 10 2025 Page 1296-1305
Francis Chigozie, Odikwa Ndubuisi, Onyemaobi Bethram Chibuzo "A Survey Comparing Specialized Hardware and Evolution in CPU, GPU AND TPU for Neural Network" Iconic Research And Engineering Journals, vol. 8, no. 10, Apr. 2025
Francis Chigozie, Odikwa Ndubuisi, Onyemaobi Bethram Chibuzo (2025). A Survey Comparing Specialized Hardware and Evolution in CPU, GPU AND TPU for Neural Network. Iconic Research And Engineering Journals, 8(10).
Francis Chigozie, Odikwa Ndubuisi, Onyemaobi Bethram Chibuzo "A Survey Comparing Specialized Hardware and Evolution in CPU, GPU AND TPU for Neural Network" Iconic Research And Engineering Journals, vol. 8, no. 10, Apr. 2025.
@article{1707828,
      author = {Francis Chigozie, Odikwa Ndubuisi, Onyemaobi Bethram Chibuzo},
      title = {A Survey Comparing Specialized Hardware and Evolution in CPU, GPU AND TPU for Neural Network},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {8},
      number = {10},
      pages = {1296-1305},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1707828.pdf},
      abstract = {This survey paper is focused on the genesis of GPU AND TPU from first generation and their architectures. The paper compares CPUs, GPUs, and TPUs, their hardware architectures, their similarities and differences were extensively discussed. The batch of data are immensely used these days but they require more time, computation and energy. Due to the greater demand and attractive options for architects to explore, companies are continuously working to reduce training and inference response time. Pertinent to the demands and cost factors different kinds of ASICs (application specific integrated circuits) are developed and research is increased in this area. Many models of CPUs, GPUs and TPUs have been developed to support these networks and to improve training and inference phase. The hardware of CPUs and GPUs can be sold to businesses while Google offers TPU processing for everyone from the cloud. When the data is away from the computational source, it increases the overall cost and to reduce this cost companies implements memory management and caching techniques close to ALUs.},
      keywords = {GPU, TPU, CPU, Deep Learning, ALU, framework.},
      month = {April},
  }