Home / Current Issue / Paper 1718503
Dynamic LoRA Rank Selection for Parameter-Efficient Fine-Tuning Under Memory-Constrained Environments
Subject area: Science,Engineering and Technology · Area of research: Artificial Intelligence
DOI: https://doi.org/10.64388/IREV9I11-1718503
Abstract
Large Language Models (LLMs) have achieved remarkable performance across various Natural Language Processing (NLP) tasks; however, fine-tuning these models requires significant computational resources and memory. Parameter-Efficient Fine-Tuning (PEFT) techniques such as Low-Rank Adaptation (LoRA) reduce training costs by updating only a small number of parameters. Traditional LoRA approaches generally use a fixed rank value throughout training, which may lead to inefficient memory utilization and suboptimal model performance in memory-constrained environments. This paper proposes a Dynamic LoRA Rank Selection approach that adaptively adjusts the rank during fine-tuning based on memory availability, model complexity, and task requirements. The proposed method aims to improve training efficiency while maintaining model accuracy and reducing computational overhead. Experimental analysis demonstrates that dynamic rank adaptation can achieve better resource utilization and comparable performance when compared to static-rank LoRA methods. The proposed approach is especially beneficial for edge devices, low-resource systems, and environments with limited GPU memory, enabling efficient deployment of large-scale AI models with reduced hardware requirements.
Keywords
Dynamic Rank Selection, Large Language Models, Low-Rank Adaptation (LoRA), Memory-Constrained Environments, Parameter-Efficient Fine-Tuning (PEFT)
References
[1] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” in Proceedings of the International Conference on Learning Representations (ICLR), 2022.
[2] B. Lester, R. Al-Rfou, and N. Constant, “The Power of Scale for Parameter-Efficient Prompt Tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021.
[3] X. L. Li and P. Liang, “Prefix-Tuning: Optimizing Continuous Prompts for Generation,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2021.
[4] N. Ding, Y. Qin, G. Yang, F. Wei, Z. Yang, Y. Su, and J. Hu, “Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models,” arXiv preprint arXiv:2203.06904, 2022.
[5] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient Finetuning of Quantized LLMs,” in Advances in Neural Information Processing Systems (NeurIPS), 2023.
[6] J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of NAACL-HLT, 2019.
[7] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Advances in Neural Information Processing Systems (NeurIPS), 2017.
[8] Z. Liu, Y. Lin, and M. Sun, “Adaptive Low-Rank Adaptation for Parameter-Efficient Fine-Tuning,” IEEE Access, vol. 12, pp. 11234–11245, 2024.
[9] S. Han, H. Mao, and W. Dally, “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” in International Conference on Learning Representations (ICLR), 2016.
[10] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, and others, “Language Models are Few-Shot Learners,” in Advances in Neural Information Processing Systems (NeurIPS), 2020.
[11] OpenAI, “GPT-4 Technical Report,” arXiv preprint arXiv:2303.08774, 2023.
[12] C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. Liu, “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” Journal of Machine Learning Research, vol. 21, no. 140, pp. 1–67, 2020.
[13] Y. LeCun, Y. Bengio, and G. Hinton, “Deep Learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015.
How to cite this paper
@article{1718503,
author = {Siddhartha Goola, C. Akhila Krishnan},
title = {Dynamic LoRA Rank Selection for Parameter-Efficient Fine-Tuning Under Memory-Constrained Environments},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {11},
pages = {4999-5007},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1718503.pdf},
abstract = {Large Language Models (LLMs) have achieved remarkable performance across various Natural Language Processing (NLP) tasks; however, fine-tuning these models requires significant computational resources and memory. Parameter-Efficient Fine-Tuning (PEFT) techniques such as Low-Rank Adaptation (LoRA) reduce training costs by updating only a small number of parameters. Traditional LoRA approaches generally use a fixed rank value throughout training, which may lead to inefficient memory utilization and suboptimal model performance in memory-constrained environments. This paper proposes a Dynamic LoRA Rank Selection approach that adaptively adjusts the rank during fine-tuning based on memory availability, model complexity, and task requirements. The proposed method aims to improve training efficiency while maintaining model accuracy and reducing computational overhead. Experimental analysis demonstrates that dynamic rank adaptation can achieve better resource utilization and comparable performance when compared to static-rank LoRA methods. The proposed approach is especially beneficial for edge devices, low-resource systems, and environments with limited GPU memory, enabling efficient deployment of large-scale AI models with reduced hardware requirements.},
keywords = {Dynamic Rank Selection, Large Language Models, Low-Rank Adaptation (LoRA), Memory-Constrained Environments, Parameter-Efficient Fine-Tuning (PEFT)},
month = {May},
doi = {https://doi.org/10.64388/IREV9I11-1718503}
}