Loading...

Interpretability of Transformer Based Language Models in NLP Tasks

Amirzadeh Gougheri, Hamid Reza | 2025

58 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 58564 (19)
  4. University: Sharif University of Technology
  5. Department: Computer Engineering
  6. Advisor(s): Sameti, Hossein
  7. Abstract:
  8. With the rapid development of Transformer-based language models such as BERT, GPT, and Llama, a major transformation has occurred in the performance of various natural language processing tasks. These models have achieved unprecedented performance in many tasks, including machine translation, sentiment analysis, and text generation. However, the inherent complexity of these models has caused their decision making processes to remain largely a black box. This issue has underscored the urgency of research in the field of interpretability of language models, as understanding their internal mechanisms is not only essential for increasing user trust, but also for ensuring responsible, fair, and reliable behavior in high stakes applications. This research investigates the interpretability of language models from the perspective of local analysis, with a particular focus on the feature attribution method. The goal of this method is to identify the contribution of each input token to the model’s output, offering useful tools for error analysis, bias detection, and the design of human-understandable explanations. In the first part of this thesis, a case study is designed to demonstrate the effectiveness of interpretability methods and to motivate the present research. This case study examines how language models handle sentences that contain multiple cues for disambiguating the gender of the subject. The findings show that encoder-based models tend to rely more on early cues, while decoder-based models rely more on the later cues in the context. These findings can inform the design of more effective user interactions with language models. In the second part of this research, a novel method is developed for interpreting transformer models based on the encoder-decoder architecture. The proposed method is built upon a representation decomposition approach, in which the internal representation of each token is rewritten as a combination of contributions from the input tokens and previously generated tokens. To address the challenges posed by the nonlinearity of certain components (feedforward networks), suitable computational approximations are employed and shown to be effective through empirical studies. Finally, by training a machine translation model and evaluating the proposed interpretability method using the Alignment Error Rate (AER) metric on the Europarl v7 dataset, it is demonstrated that the proposed method improves AER by 1.3% over the best existing baseline.
  9. Keywords:
  10. Transformer-based Language Models ; Interpretability ; Encoder-Decoder ; Machine Translation ; Representation Decomposition ; Feature Importance ; Local Analysis

 Digital Object List

 Bookmark

...see more