Loading...
Search
Search in this resource
sort by
تشخیص فعالیت افراد در ویدیو با بهره‌گیری از انتقال دانش مدل زبانی و استفاده از ساز و کار توجه
1026 viewed

تشخیص فعالیت افراد در ویدیو با بهره‌گیری از انتقال دانش مدل زبانی و استفاده از ساز و کار توجه

ابوالقاسمی، مرتضی Abolghasemi,Morteza

Video Action Recognition Using Transfer Learning of Language Model and Attention Mechanism

Abolghasemi,Morteza | 2023

226 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 57319 (19)
  4. University: Sharif University of Technology
  5. Department: Computer Engineering
  6. Advisor(s): Rabiee, Hamid Reza
  7. Abstract:
  8. In the new era of machine vision, action recognition in videos remains a pivotal and essential challenge. With advancements in processing capabilities and the rise of digital data, multimodal networks like CLIP (Contrastive Language-Image Pre-training) have emerged, adeptly bridging the connection between visual and linguistic data. While these networks are tailor-made for images and associated texts, they falter when confronted with motion data in videos. In this research, we propose a novel dual-stream architecture aiming to augment action recognition in videos, astutely blending the encoding prowess of the VideoMAE model with the representational strength of the CLIP model. Our design employs the VideoMAE encoder to process multiple contiguous frames, capturing both the inherent content and interrelations. Subsequently, a vision transformer layer facilitates the seamless transfer of embeddings from the VideoMAE domain to the CLIP space. Our dual-stream approach ensures the optimal extraction of both dynamic and static video content. Additionally, we have enriched training and evaluation labels using the GPT-4 large language model, converting them into descriptive video sentences. Contrastive learning techniques form the bedrock of our training objective. Evaluations on the UCF101 dataset manifest the model's efficacy, indicating the proficient extraction of video motion information via the VideoMAE network—especially notable due to its Masked Auto-Encoders training on a colossal UnlabeledHybrid video dataset, positioning it as one of the video foundation models.) Ultimately, we have achieved a successful transfer of dynamic video data into the shared CLIP text-image space
  9. Keywords:
  10. Action Recognition ; Transfer Learning ; Attention Mechanism ; Contrastive Learning ; Pretrained Models ; Zero-Shot Learning ; Multimodal Networks

 Digital Object List

 Bookmark

  • فصل مقدمه
    • اهداف پژوهش
  • فصل کارهای پیشین
    • روش‌های اولیه و مسیر‌های متراکم
    • شبکه‌های عصبی پیچشی
    • پویایی زمانی با کمک شبکه‌های بازگشتی
    • شبکه‌های پیچشی سه بعدی
    • تشخیص کنش مبتنی بر اسکلت
    • مکانیسم توجه و شبکه‌های مبدلی در ویدیوها
    • یادگیری چندوجهی
    • تکنیک‌های خود نظارت و کدگذاری پیشرفته
  • فصل راه کار پیشنهادی
    • ایده اولیه
    • معماری مدل
      • کدگذار
      • لایه‌ی مبدلی
      • رویکرد دو جریانی
    • دادگان
      • تقویت برچسب
  • فصل جزئیات پیاده‌سازی و نتایج
    • آموزش با کمک یادگیری متضاد
    • تنظیمات پیش‌آموزش
    • یادگیری پس از پیش‌آموزش
    • ارزیابی
    • تجزیه و تحلیل پیکره‌بندی‌های مختلف
    • مشارکت اصلی
  • فصل جمع بندی
    • نتیجه گیری
    • کار‌های آینده
      • واحد انتخاب قاب
      • گسترش به مجموعه داده‌های متنوع تر
      • ادغام ورودی‌های چندوجهی
      • تنظیم دقیق پارامتر‌ها
  • پیوست
  • مراجع
  • واژه نامه انگلیسی به فارسی
  • واژه نامه فارسی به انگلیسی
...see more