Loading...

Stochastic fine-tuning of language models using masked gradients

Akbar-Tajari, M | 2024

263 Viewed
  1. Type of Document: Article
  2. DOI: 10.18653/v1/2024.findings-emnlp.1002
  3. Publisher: ACL Anthology , 2024
  4. Abstract:
  5. Large Language Models (LLMs) have emerged as the dominant paradigm in Natural Language Processing owing to their remarkable performance across various target tasks. However, naively fine-tuning them for specific downstream tasks often requires updating a vast number of parameters, resulting in high computational costs and overfitting when training data is limited. In this paper, we propose a novel approach, called Stochastic Tuning, that addresses these challenges by selectively updating a small subset of parameters in each step of the tuning process. Our approach is characterized by its customization of updates based on task-specific partial gradients with respect to stochastic sub-networks. The advantage of Stochastic Tuning over existing solutions lies in its ability to consider both parameter weights as well as forward values which guarantees a context-sensitive fine-tuning. Our experiments demonstrate that Stochastic Tuning outperforms existing lightweight fine-tuning methods, improving average performance by over two points on RoBERTa across several tasks in the GLUE benchmark while updating merely 0.08% of the model's parameters. The code for our implementation can be found at https://github.com/m-Tajari/StocTuning_LLMs. © 2024 Association for Computational Linguistics
  6. Keywords:
  7. Benchmarking ; Computational linguistics ; Computer circuits ; Context sensitive languages ; Gluing ; Modeling languages ; Natural language processing systems ; Stochastic models ; Computational costs
  8. Source: EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024 ; 2024 , Pages 17195-17202 ; 979-889176168-1 (ISBN)
  9. URL: https://aclanthology.org/2024.findings-emnlp.1002.pdf