Loading...
- Type of Document: M.Sc. Thesis
- Language: Farsi
- Document No: 58280 (19)
- University: Sharif University of Technology
- Department: Computer Engineering
- Advisor(s): Kasaei, Shohreh
- Abstract:
- Accurate classification of chest radiography images and retrieval of relevant medical reports have become significant challenges in the medical field due to the shortage of radiology specialists and the continuous increase in medical image volumes. The importance of this issue is highlighted by the direct impact that delays in diagnosing cardiac and pulmonary diseases have on public health. Key applications of this area include automatic and precise disease classification, accelerated generation of medical reports, and the development of intelligent systems to support clinical decision-making, especially in resource-limited areas. Existing methods primarily utilize either image or text data independently, resulting in significant limitations when dealing with multi-label data, complexities of image and text retrieval, and the inability to establish meaningful semantic relationships between different modalities. This research introduces a novel approach based on the analysis and fusion of multimodal data. The proposed method employs two separate encoders for extracting visual and textual features, which are semantically aligned and integrated through attention mechanisms and contrastive learning. The main innovations of the proposed method include the design of intra-modal and inter-modal contrastive loss functions, the use of the Jaccard similarity criterion for intelligent data sampling, and a two-stage attention structure to enhance the extracted features. This process simultaneously improves the performance in classifying 14 types of pulmonary diseases and retrieving relevant medical reports. Limitations of this research include the scarcity of paired image-text data and the high variability in report structures. To evaluate the proposed approach, two standard and widely recognized datasets, MIMIC-CXR and CheXpert, were employed. The performance of the proposed method was assessed using precise retrieval metrics such as Precision@k and Rank@k. The results demonstrated that the proposed method achieved a classification accuracy of 94.9%, a 14% improvement in lower values of k in Precision@k, and the highest successful retrieval rate across all values of k in Rank@k, indicating a significant enhancement over previous methods. Consequently, the proposed method can serve as an effective tool for supporting clinical diagnosis and decision-making processes
- Keywords:
- Medical Images ; Contrastive Learning ; Vision-Language Models ; Multimodal Learning ; Images Classification ; MIMIC-CXR Dataset
-
محتواي کتاب
- view
