Loading...
ViT-PMN: A vision transformer approach for persian numeral recognition
Ardehkhani, P ; Sharif University of Technology | 2024
43
Viewed
- Type of Document: Article
- DOI: 10.1109/AISP61396.2024.10475234
- Publisher: 2024
- Abstract:
- This study focuses on the task of Persian numeral classification within image data, employing the Vision Transformer (ViT) architecture to predict numerals akin to the MNIST dataset, but adapted to the Persian script. Our approach yielded a notable validation accuracy of 0.9920, particularly noteworthy when employing a patch size of 4. Notably, the research introduces an innovative visualization aspect, showcasing the first multi-head attention linear map and its counterpart, the last one. The visualization of these attention maps provides a unique insight into the model's internal processes and highlights its proficiency in capturing intricate patterns within Persian numeral images. This work contributes to the evolving landscape of character recognition, specifically addressing the challenges posed by the Persian script, and underscores the efficacy of employing the ViT architecture for such intricate tasks. The achieved validation accuracy and the detailed visualization of attention maps mark notable milestones in the realm of Persian numeral classification, showcasing the potential of Vision Transformers in the context of script-specific optical character recognition. © 2024 IEEE
- Keywords:
- Artificial Intelligence ; Classification ; Persian Dataset ; Vision Transformer ; Classification (of information) ; Optical character recognition ; Deep learning ; Image data ; Linear maps ; Numeral recognition ; Patch size ; Persians ; Visualization
- Source: 2024 20th CSI International Symposium on Artificial Intelligence and Signal Processing, AISP 2024 ; 2024 ; 979-835038394-2 (ISBN)
- URL: https://ieeexplore.ieee.org/document/10475234
