Loading...

Some Model-free Discrete Reinforcement Learning Algorithms

Yousefizadeh, Hossein | 2021

1251 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 54227 (02)
  4. University: Sharif University of Technology
  5. Department: Mathematical Sciences
  6. Advisor(s): Daneshgar, Amir
  7. Abstract:
  8. In this thesis, we review some methods related to model-free discrete reinforcement learning and their corresponding algorithms. Our main goal is to present existing methods in an integrated and formal setup, without compromising their mathematical accuracy or comprehensibility. We have done our best to fix the inconsistencies existing in notations and definitions appearing in different areas of the vast literature. We discuss dynamic programming methods, including policy iteration and value iteration and temporal difference methods as well as policy-based methods such as policy gradient, advantage actor-critic, TRPO, and PPO. Among value-based methods, we discuss Q-learning and C51 where we also review some intermediate methods which use the ideas of both approaches, such as DDPG And SAC. We use examples when necessary to clarify the concepts and methods. Finally, to summarize, we provide some comparative analysis of these methods discussed
  9. Keywords:
  10. Reinforcement Learning ; Deep Learning ; Machine Learning ; Dynamic Programming ; Discrete Reinforcement Learning

 Digital Object List

 Bookmark

  • 0a25dc607c6f278af7d03a356b89bc029cbea1b8cedc3fda80cc678446349530.pdf
  • bb8933c2089708c26f8d4506b7aebc7f2c7f5a619f5c8ff25960e8272a2ecd4a.pdf
    • جمع بندی
    • مقایسه
    • مقایسه و جمع بندی
        • روش SAC
        • روش DDPG
      • روش‌های میانی
        • روش DQN
      • روش‌های مبتنی بر ارزش
        • روش PPO
        • روش TRPO
        • روش‌های بازیگر-منتقد
        • الگوریتم REINFORCE
        • روش گرادیان خط‌مشی
      • روش‌های مبتنی بر خط‌مشی
      • دو رویکرد مختلف به مسئله
    • معرفی چند الگوریتم جدید
        • روش های Q-learning
      • روش های یادگیری تفاوت زمانی
        • الگوریتم های تکرار خط‌مشی تعمیم یافته
        • الگوریتم تکرار ارزش
        • الگوریتم تکرار خط‌مشی
        • بهبود خط‌مشی
        • برنامه‌ریزی پویا
        • بهینگی و معادله بهینگی بلمن
        • معادله بلمن
        • خط‌مشی و تابع ارزش بهینه
      • برنامه‌ریزی پویا
        • عایدی و تابع ارزش
        • خط‌مشی
      • نگاهی دقیق‌تر به عناصر اصلی یادگیری تقویتی
      • یک مثال ساده: ربات قوطی یاب
      • فرایند تصمیم‌گیری مارکوف
    • مفاهیم اولیه و الگوریتم‌های کلاسیک
        • مدل محیط
        • تابع ارزش
        • سیگنال پاداش
        • خط‌مشی
      • عناصر اصلی یادگیری تقویتی
        • کارهای دوره‌ای و مستمر
        • اکتشاف و بهره‌برداری
        • دینامیک عامل-محیط
        • آینده نگری و برنامه ریزی
      • یادگیری تقویتی چیست
      • تاریخچه یادگیری تقویتی
      • دسته‌بندی روش‌های یادگیری ماشین
      • ساختار پایان‌نامه
    • مقدمه
...see more