Bayes Politika Arama için Markov Zinciri Monte Carlo Algoritması
2019
0 views
0 downloads
Advisor: Assoc. Prof. Dr. Ahmet Onat ; Dr. Sinan Yıldırım
Abstract (EN)
The fundamental intention in Reinforcement Learning (RL) is to seek for optimal parameters of a given parameterized policy. Policy search algorithms have paved the way for making the RL suitable for applying to complex dynamical systems, such as the robotics domain, where the environment comprised of high-dimensional state and action spaces. Although many policy search techniques are based on the widespread policy gradient methods, thanks to their appropriateness to such complex environments, their performance might be affected by slow convergence or local optima complications. The reason for this is due to the urge for computation of the gradient components of the parameterized policy. In this study, we avail a Bayesian approach for policy search problem pertinent to the RL framework, The problem of interest is to control a discrete-time Markov decision process (MDP) with continuous state and action spaces. We contribute to the field by propounding a Particle Markov Chain Monte Carlo (P-MCMC) algorithm as a method of generating samples for the policy parameters from a posterior distribution, instead of performing gradient approximations. To do so, we adopt a prior density over policy parameters and aim for the posterior distribution where the 'likelihood' is assumed to be the expected total reward. In terms of risk-sensitive scenarios, where a multiplicative expected total reward is employed to measure the performance of the policy, rather than its cumulative counterpart, our methodology is fit for purpose owing to the fact that by utilizing a reward function in a multiplicative form, one can fully take sequential Monte Carlo (SMC), known as the particle filter within the iterations of the P-MCMC. it is worth mentioning that these methods have widely been used in statistical and engineering applications in recent years. Furthermore, in order to
Author
Dr. Vahıd Tavakol Aghaeı
How to Cite
Vahıd Tavakol Aghaeı (Doctorate thesis). Bayes Politika Arama için Markov Zinciri Monte Carlo Algoritması, 2019, Sabanci University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Sabanci University
- Popülizm, bozulmalar ve kriz algısı(2019)
- Görme biçimleri: Nevizâde Atai'nin Alemnüma'sı ve 17. yüzyılın başlarında Osmanlı toplumunun görsel algısında değişimler(2020)
- Kim Var Orada? çağdaş Türkiye tiyatrosu'nda sessizleştirilmiş geçmişleri sahnelemek: Kim Var Orada? Muhsin Bey'in Son Hamleti(2020)
- İstanbul'da bulunan fahişelerin Geç Osmanlı Dönemi'ndeki yaşamlarının Ahmed Midhat Efendi ve Hüseyin Rahmi Gürpınar romanları üzerinden bir değerlendirmesi(2019)
- Sınırların yeniden çizilmesi: Üniversite öğrencilerinin sözlü tarihi(2020)
- Normal ve genelleştirilmiş bir gamma popülasyonundaki m'inci (merkezi) moment için maksimum olabilirlik ve örnek momenti tahmin edicisi üzerine(2020)
