Loading...
Search
| Friend's email | |
| Your name | |
| Your email | |
| enter code | |
This page was sent successfuly
55 viewed
بررسی روشمند و بهبود در آموزش و استنتاج مدل های زبانی با محدودیت منابع سخت افزاری
پاک سرشت، محمد سینا Pakseresht, Mohammad Sina
Optimizing Language Model Training and Inference: A Systematic Approach for Resource-Constrained Conditions
Pakseresht, Mohammad Sina | 2025
55
Viewed
- Type of Document: M.Sc. Thesis
- Language: Farsi
- Document No: 58655 (19)
- University: Sharif University of Technology
- Department: Computer Engineering
- Advisor(s): Asgari, Ehsaneddin; Soleymani Baghshah, Mahdieh
- Abstract:
- In recent years, Large Language Models have demonstrated remarkable capabilities in solving a wide range of natural language processing tasks. However, their high computational cost has made their use in resource-constrained environments a significant challenge. Given these challenges, Small Language Models (SLMs) have emerged as an efficient alternative, especially for specialized applications under severe hardware constraints. This thesis proposes a novel framework for structured, dynamic, and task-specific pruning of SLM layers, aimed at reducing computational cost and inference latency, specifically optimized for the tool-calling task. The proposed method introduces a multi-stage pipeline. First, instead of permanently removing Transformer layers, they are replaced with lightweight Low-Rank Adaptation adapters. These adapters are trained individually using white-box knowledge distillation to mimic the hidden-state transformation of their original layer with a negligible fraction of the parameters. To enable dynamic pruning, a small, BERT-based router model is introduced, which decides which layers and how many of them to prune for each input by assessing its complexity. This router is trained via offline reinforcement learning, leveraging preference-based algorithms like Direct Preference Optimization (DPO) and Odds Ratio Preference Optimization (ORPO). This process is performed on a pre-generated offline dataset of various pruning configurations and their corresponding performance scores, which eliminates the need for costly interaction with the environment during training. This framework is implemented on the 1-billion-parameter xLAM-2 model and a specialized tool-calling dataset. The results indicate that this method can increase the model's speed by approximately 18\% by dynamically removing layers, with an accompanying accuracy drop of about 10\%. This decrease in accuracy is comparable to statically (non-dynamically) removing 3 layers; however, the dynamic approach achieved the removal of 6 layers from the model
- Keywords:
- Resource Constraint ; Low-Rank Adaptation (LoRA) ; Tool Calling ; Dynamic Pruning ; Structured Pruning ; White-Box Knowledge Distillation ; Offline Reinforcement Learning ; Small Language Models (SLMs)
-
محتواي کتاب
- view
