Home / Current Issue / Paper 1706957
Transformers Beyond NLP: Expanding Horizons in Machine Learning
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
Abstract
Transformers, initially designed for natural language processing (NLP), have revolutionized machine learning with their self-attention mechanisms and unparalleled scalability. Originally developed for tasks such as machine translation and text summarization, transformers have demonstrated exceptional performance in capturing complex dependencies and contextual relationships within sequential data. Their success in NLP has inspired researchers to adapt these architectures for various other domains. By leveraging the unique properties of self-attention and multi-head attention, transformers have been reimagined to process visual data, model temporal patterns, and analyze biological sequences with remarkable accuracy and efficiency. Furthermore, their application in generative modeling has paved the way for innovations in creative AI, including text-to-image synthesis and music composition. This paper provides a comprehensive overview of how transformers have transcended their initial domain, driving advancements in fields as diverse as computer vision, bioinformatics, time-series analysis, and beyond. Challenges such as computational demands, data requirements, and interpretability are also discussed, along with future directions to address these limitations and expand their transformative potential.
Keywords
Transformers, Self-Attention, Machine Learning, Neural Networks, Computer Vision, Bioinformatics, Time- Series Analysis, Generative Modeling, Efficient Architectures, Artificial Intelligence, Cross-Modal Learning, Interpretability, Scalability, Sustainability, Domain-Specific Applications
References
[1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, pp. 5998–6008, 2017.
[2] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021.
[3] N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. Eur. Conf. Comput. Vis. (ECCV), pp. 213–229, 2020.
[4] J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, et al., “Highly accurate protein structure prediction with AlphaFold,” Nature, vol. 596, no. 7873, pp. 583–589, 2021.
[5] Y. Ji, Z. Zhou, H. Liu, Y. Wang, and J. Zheng, “DNABERT: Pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome,” Bioinformatics, vol. 37, no. 15, pp. 2112– 2120, 2021.
[6] B. Lim, S. Arik, N. Loeff, and T. Pfister, “Temporal fusion transformers for interpretable multi-horizon time series forecasting,” Int. J. Forecast- ing, vol. 37, no. 4, pp. 1748–1764, 2021.
[7] H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang, “Informer: Beyond efficient transformer for long sequence time-series forecasting,” in Proc. Assoc. Advancement Artif. Intell. (AAAI), vol. 35, no. 12, pp. 11106–11115, 2021.
[8] A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in Proc. Int. Conf. Mach. Learn. (ICML), vol. 139, pp. 8821–8831, 2021.
[9] R. Child, S. Gray, A. Radford, and I. Sutskever, “Generating long sequences with sparse transformers,” arXiv preprint arXiv:1904.10509, 2019.
[10] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. Assoc. Comput. Linguistics, pp. 4171–4186, 2019.
[11] S. Wang, B. Li, M. Khabsa, H. Fang, and H. Ma, “Linformer: Self- attention with linear complexity,” arXiv preprint arXiv:2006.04768, 2020.
[12] K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins et al., “Rethinking attention with performers,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021.
[13] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Je´gou, “Training data-efficient image transformers & distillation through attention,” in Proc. Int. Conf. Mach. Learn. (ICML), vol. 139, pp. 10347– 10357, 2021.
[14] S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, P. Fu, and L. Zhang, “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 6881–6890, 2021.
[15] Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, and H. Tong, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 10012–10022, 2021.
[16] L. Yuan, Y. Chen, T. Wang, W. Yu, Z. Shi, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token ViT: Training vision transformers from scratch on ImageNet,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 558–567, 2021.
[17] B. Fabian, A. Edinger, M. Filip, K. Claudia, B. Tim, and R. Martin, “Molecular property prediction: A transformer-based architecture for modeling molecular graphs,” arXiv preprint arXiv:2011.07457, 2020.
[18] Z. Avsec, J. Agarwal, B. Visentin, D. Ledsam, A. Grabska-Barwinska, J. Taylor, and D. Kelley, “Effective gene expression prediction from sequence by integrating long-range interactions,” Nature Methods, vol. 18, no. 10, pp. 1196–1203, 2021.
[19] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, pp. 5998–6008, 2017.
[20] B. Lim, S. Arik, N. Loeff, and T. Pfister, “Temporal fusion transformers for interpretable multi-horizon time series forecasting,” Int. J. Forecast- ing, vol. 37, no. 4, pp. 1748–1764, 2021.
[21] H. Wu, J. Xu, J. Wang, and F. Long, “Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting,” in Advances in Neural Information Processing Systems, vol. 34, pp. 22419– 22430, 2021. bibitemradford2019language A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI, 2019.
[22] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, Neelakantan et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems, vol. 33, pp. 1877– 1901, 2020.
[23] C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou et al., “Exploring the limits of transfer learning with a unified text- to-text transformer,” J. Mach. Learn. Res., vol. 21, no. 140, pp. 1–67, 2020.
[24] M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence- to-sequence pretraining for natural language generation, translation, and comprehension,” in Proc. Assoc. Comput. Linguistics, pp. 7871–7880, 2020.
[25] C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. Gontijo- Lopes et al., “Imagen: Text-to-image diffusion models,” arXiv preprint arXiv:2205.11487, 2022.
[26] P. Dhariwal, H. Jun, C. Payne, J. Kim, A. Radford, and I. Sutskever, “Jukebox: A generative model for music,” OpenAI, 2020.
[27] X. Yan, J. Xu, X. Dai, and X. Zhou, “VideoGPT: Generative pretraining for videos,” arXiv preprint arXiv:2104.10157, 2021.
[28] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,” in Proc. Eur. Conf. Comput. Vis. (ECCV), pp. 405–421, 2020.
[29] Y. Jiang, S. Zhang, W. Gong, X. Zheng, and Z. Li, “TransGAN: Two transformers can make one strong GAN,” in Proc. Adv. Neural Inf. Process. Syst., vol. 34, pp. 14742–14754, 2021.
[30] R. Child, S. Gray, A. Radford, and I. Sutskever, “Generating long sequences with sparse transformers,” arXiv preprint arXiv:1904.10509, 2019.
[31] S. Wang, B. Li, M. Khabsa, H. Fang, and H. Ma, “Linformer: Self- attention with linear complexity,” arXiv preprint arXiv:2006.04768, 2020.
[32] K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins et al., “Rethinking attention with performers,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021.
[33] M. Zaheer, G. Guruganesh, K. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, and A. Vaswani, “Big bird: Transformers for longer sequences,” in Advances in Neural Information Processing Systems, vol. 33, pp. 17283–17297, 2020.
[34] A. Jaegle, S. Gimeno, S. Brockman, L. Zong, C. Voss, J. Lapedriza, L. Kaplan et al., “Perceiver: General perception with iterative attention,” in Proc. Int. Conf. Mach. Learn. (ICML), vol. 139, pp. 4651–4663, 2021.
[35] A. Radford, J. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, and P. Dhariwal, “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn. (ICML), vol. 139, pp. 8748–8763, 2021.
[36] E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy consid- erations for deep learning in NLP,” in Proc. Assoc. Comput. Linguistics, pp. 3645–3650, 2019.
[37] E. Ganesh, J. Perez, M. Ranzato, and D. Grangier, “Compressing transformers with low-rank and sparse approximations,” arXiv preprint arXiv:2112.05682, 2021.
[38] L. Yuan, J. Chen, C. Wang, Z. Wang, Y. Feng, Z. Shen, C. Guo et al., “Florence: A new foundation model for computer vision,” arXiv preprint arXiv:2111.11432, 2021.
[39] S. Jain and B. C. Wallace, “Attention is not explanation,” in Proc. Assoc. Comput. Linguistics, pp. 3543–3556, 2019.
How to cite this paper
@article{1706957,
author = {Srikanth Kamatala, Anil Kumar Jonnalagadda, Prudhvi Naayini},
title = {Transformers Beyond NLP: Expanding Horizons in Machine Learning},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {8},
number = {7},
pages = {441-452},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1706957.pdf},
abstract = {Transformers, initially designed for natural language processing (NLP), have revolutionized machine learning with their self-attention mechanisms and unparalleled scalability. Originally developed for tasks such as machine translation and text summarization, transformers have demonstrated exceptional performance in capturing complex dependencies and contextual relationships within sequential data. Their success in NLP has inspired researchers to adapt these architectures for various other domains. By leveraging the unique properties of self-attention and multi-head attention, transformers have been reimagined to process visual data, model temporal patterns, and analyze biological sequences with remarkable accuracy and efficiency. Furthermore, their application in generative modeling has paved the way for innovations in creative AI, including text-to-image synthesis and music composition. This paper provides a comprehensive overview of how transformers have transcended their initial domain, driving advancements in fields as diverse as computer vision, bioinformatics, time-series analysis, and beyond. Challenges such as computational demands, data requirements, and interpretability are also discussed, along with future directions to address these limitations and expand their transformative potential.},
keywords = {Transformers, Self-Attention, Machine Learning, Neural Networks, Computer Vision, Bioinformatics, Time- Series Analysis, Generative Modeling, Efficient Architectures, Artificial Intelligence, Cross-Modal Learning, Interpretability, Scalability, Sustainability, Domain-Specific Applications},
month = {January},
}