Loading...
- Type of Document: M.Sc. Thesis
- Language: Farsi
- Document No: 54227 (02)
- University: Sharif University of Technology
- Department: Mathematical Sciences
- Advisor(s): Daneshgar, Amir
- Abstract:
- In this thesis, we review some methods related to model-free discrete reinforcement learning and their corresponding algorithms. Our main goal is to present existing methods in an integrated and formal setup, without compromising their mathematical accuracy or comprehensibility. We have done our best to fix the inconsistencies existing in notations and definitions appearing in different areas of the vast literature. We discuss dynamic programming methods, including policy iteration and value iteration and temporal difference methods as well as policy-based methods such as policy gradient, advantage actor-critic, TRPO, and PPO. Among value-based methods, we discuss Q-learning and C51 where we also review some intermediate methods which use the ideas of both approaches, such as DDPG And SAC. We use examples when necessary to clarify the concepts and methods. Finally, to summarize, we provide some comparative analysis of these methods discussed
- Keywords:
- Reinforcement Learning ; Deep Learning ; Machine Learning ; Dynamic Programming ; Discrete Reinforcement Learning
-
محتواي کتاب
- view
- 0a25dc607c6f278af7d03a356b89bc029cbea1b8cedc3fda80cc678446349530.pdf
- bb8933c2089708c26f8d4506b7aebc7f2c7f5a619f5c8ff25960e8272a2ecd4a.pdf
- جمع بندی
- مقایسه
- مقایسه و جمع بندی
- روش SAC
- روش DDPG
- روشهای میانی
- روش DQN
- روشهای مبتنی بر ارزش
- روش PPO
- روش TRPO
- روشهای بازیگر-منتقد
- الگوریتم REINFORCE
- روش گرادیان خطمشی
- روشهای مبتنی بر خطمشی
- دو رویکرد مختلف به مسئله
- معرفی چند الگوریتم جدید
- روش های Q-learning
- روش های یادگیری تفاوت زمانی
- الگوریتم های تکرار خطمشی تعمیم یافته
- الگوریتم تکرار ارزش
- الگوریتم تکرار خطمشی
- بهبود خطمشی
- برنامهریزی پویا
- بهینگی و معادله بهینگی بلمن
- معادله بلمن
- خطمشی و تابع ارزش بهینه
- برنامهریزی پویا
- عایدی و تابع ارزش
- خطمشی
- نگاهی دقیقتر به عناصر اصلی یادگیری تقویتی
- یک مثال ساده: ربات قوطی یاب
- فرایند تصمیمگیری مارکوف
- مفاهیم اولیه و الگوریتمهای کلاسیک
- مدل محیط
- تابع ارزش
- سیگنال پاداش
- خطمشی
- عناصر اصلی یادگیری تقویتی
- کارهای دورهای و مستمر
- اکتشاف و بهرهبرداری
- دینامیک عامل-محیط
- آینده نگری و برنامه ریزی
- یادگیری تقویتی چیست
- تاریخچه یادگیری تقویتی
- دستهبندی روشهای یادگیری ماشین
- ساختار پایاننامه
- مقدمه
