Loading...
Video Action Recognition Using Transfer Learning of Language Model and Attention Mechanism
Abolghasemi,Morteza | 2023
226
Viewed
- Type of Document: M.Sc. Thesis
- Language: Farsi
- Document No: 57319 (19)
- University: Sharif University of Technology
- Department: Computer Engineering
- Advisor(s): Rabiee, Hamid Reza
- Abstract:
- In the new era of machine vision, action recognition in videos remains a pivotal and essential challenge. With advancements in processing capabilities and the rise of digital data, multimodal networks like CLIP (Contrastive Language-Image Pre-training) have emerged, adeptly bridging the connection between visual and linguistic data. While these networks are tailor-made for images and associated texts, they falter when confronted with motion data in videos. In this research, we propose a novel dual-stream architecture aiming to augment action recognition in videos, astutely blending the encoding prowess of the VideoMAE model with the representational strength of the CLIP model. Our design employs the VideoMAE encoder to process multiple contiguous frames, capturing both the inherent content and interrelations. Subsequently, a vision transformer layer facilitates the seamless transfer of embeddings from the VideoMAE domain to the CLIP space. Our dual-stream approach ensures the optimal extraction of both dynamic and static video content. Additionally, we have enriched training and evaluation labels using the GPT-4 large language model, converting them into descriptive video sentences. Contrastive learning techniques form the bedrock of our training objective. Evaluations on the UCF101 dataset manifest the model's efficacy, indicating the proficient extraction of video motion information via the VideoMAE network—especially notable due to its Masked Auto-Encoders training on a colossal UnlabeledHybrid video dataset, positioning it as one of the video foundation models.) Ultimately, we have achieved a successful transfer of dynamic video data into the shared CLIP text-image space
- Keywords:
- Action Recognition ; Transfer Learning ; Attention Mechanism ; Contrastive Learning ; Pretrained Models ; Zero-Shot Learning ; Multimodal Networks
