Loading...
Search
Search in this resource
sort by
شناسایی موجودیت‌های نامدار در زبان فارسی با استفاده از یادگیری ژرف
405 viewed

شناسایی موجودیت‌های نامدار در زبان فارسی با استفاده از یادگیری ژرف

صبحی، محمد Sobhi, Mohamad

Named Entity Recognition in Persian Using Deep Learning

Sobhi, Mohamad | 2021

405 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 54727 (19)
  4. University: Sharif University of Technology
  5. Department: Computer Engineering
  6. Advisor(s): Sameti, Hossein
  7. Abstract:
  8. Named Entity Recognition (NER) is a key component and the first step of many natural language processing tasks such as question answering systems, information retrieval, machine translation, text summarization, and so on. First NER system initially used rule-based and machine learning methods, which grew significantly with the advent of deep learning architectures as well as the development of hardware and data resources. Traditional deep learning methods used convolutional and recursive neural networks that had disadvantages such as gradient vanishing and non-parallel computing, respectively. In addition, the need for huge corpus and powerful hardware resources was one of the problems of training these networks. In recent years, with the advent of Transformers and Transformer-based language models along with the transfer learning movement, a new era in natural language processing has emerged.In this study, we developed a NER system with the help of sequential transfer learning of pre-trained language models and its adaptation to the target task. In our proposed solution, the language model has three features based on Transformer, Masked language models and fine-tuning on a Persian corpus. Hence, in the proposed architecture, we used three different language models: multilingual BERT, XL-RoBERTa and ParsBERT as the input representation and content encoder, and a fully connected neural network as the tag decoder. In the last stage, these three models were trained on our labeled corpus, which were able to obtain F-score 84.41, 82.23 and 91.15 on Arman Corpus and 90.54, 90.2 and 97.61 on PEYMA Corpus, respectively. The comparison of evaluation results on these three language models helped us identify the most important parameters for improving NER tasks, including that monolingual language models are better than multilingual languages for NER tasks
  9. Keywords:
  10. Natural Language Processing ; Named Entity Recognition ; Transformer-based Language Models ; Sequential Transfer Learning ; Deep Learning

 Digital Object List

 Bookmark

No TOC