Loading...

Multilingual Multimodal Models for Information Retrieval

Aminian Nedushan, Arman | 2026

34 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 58813 (19)
  4. University: Sharif University of Technology
  5. Department: Computer Engineering
  6. Advisor(s): Asgari, Ehsaneddin; Sharifi Zarchi, Ali
  7. Abstract:
  8. This study explores the problem of multimodal information retrieval in English and Persian, with a particular focus on electronic commerce where users often search for products using textual or visual queries. To enable multimodal retrieval in Persian, a Persian CLIP model (Contrastive Language-Image Pretraining) was trained. The training dataset was constructed from multiple sources, including machine-translated English corpora, web-crawled Persian data from the Divar platform, and user search and click logs from the Torob marketplace. After filtering and quality evaluation, the model successfully learned a shared embedding space between Persian text and images, forming the foundation for a practical multimodal retrieval system. In the next stage, an intelligent multi-agent system was developed to combine similarity-based retrieval with analytical reasoning, enabling natural, conversational interaction with users. Experimental results show that the Persian CLIP model outperforms the original English CLIP (ACC@1 = 0.25 vs. 0.04), and the multi-agent system achieves the best overall performance with ACC@1 = 0.78. These findings demonstrate that combining multimodal models with intelligent agents can significantly enhance user experience and retrieval accuracy in Persian-language search systems
  9. Keywords:
  10. Multimodal Information Retrieval ; Electronic Commerce ; Intelligent Multi-Agent System ; Contrastive Language–Image Pretraining (CLIP) Model ; Localized Dataset

 Digital Object List

 Bookmark

...see more