Bayes Politika Arama için Markov Zinciri Monte Carlo Algoritması
Bu tez size mi ait?
Bu kayıt toplu arşivden geldi. Sizinse profilinize bağlayın.
2019
0 görüntülenme
0 i̇ndirme
Danışman: Assoc. Prof. Dr. Ahmet Onat ; Dr. Sinan Yıldırım
Özet (EN)
The fundamental intention in Reinforcement Learning (RL) is to seek for optimal parameters of a given parameterized policy. Policy search algorithms have paved the way for making the RL suitable for applying to complex dynamical systems, such as the robotics domain, where the environment comprised of high-dimensional state and action spaces. Although many policy search techniques are based on the widespread policy gradient methods, thanks to their appropriateness to such complex environments, their performance might be affected by slow convergence or local optima complications. The reason for this is due to the urge for computation of the gradient components of the parameterized policy. In this study, we avail a Bayesian approach for policy search problem pertinent to the RL framework, The problem of interest is to control a discrete-time Markov decision process (MDP) with continuous state and action spaces. We contribute to the field by propounding a Particle Markov Chain Monte Carlo (P-MCMC) algorithm as a method of generating samples for the policy parameters from a posterior distribution, instead of performing gradient approximations. To do so, we adopt a prior density over policy parameters and aim for the posterior distribution where the 'likelihood' is assumed to be the expected total reward. In terms of risk-sensitive scenarios, where a multiplicative expected total reward is employed to measure the performance of the policy, rather than its cumulative counterpart, our methodology is fit for purpose owing to the fact that by utilizing a reward function in a multiplicative form, one can fully take sequential Monte Carlo (SMC), known as the particle filter within the iterations of the P-MCMC. it is worth mentioning that these methods have widely been used in statistical and engineering applications in recent years. Furthermore, in order to
Yazar
Vahıd Tavakol Aghaeı
Bu Yayına Nasıl Atıf Yapılır
Vahıd Tavakol Aghaeı (Doctorate thesis). Bayes Politika Arama için Markov Zinciri Monte Carlo Algoritması, 2019, Sabancı University.
Anahtar Kelimeler
Lisans
Tüm Hakları Saklıdır
Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.
Sabancı University tezlerinden daha fazlası
- Türkiye'de parti içi demokrasi: Teori, görünüm ve sorunlar(2019)
- Gelinin bedeli / berdel ve çocuk evliliklerinin görsel anlatısı(2019)
- Popülizm, bozulmalar ve kriz algısı(2019)
- Görme biçimleri: Nevizâde Atai'nin Alemnüma'sı ve 17. yüzyılın başlarında Osmanlı toplumunun görsel algısında değişimler(2020)
- Kim Var Orada? çağdaş Türkiye tiyatrosu'nda sessizleştirilmiş geçmişleri sahnelemek: Kim Var Orada? Muhsin Bey'in Son Hamleti(2020)
- Bir mahallede iki dünyanın buluşması: İstanbul, Yenimahalle'de göç alan toplum üyeleri ve Afgan işçilerin ilişkileri(2019)
