Keyframe demonstration seeded and Bayesian optimized policy search
2022
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Barış Akgün
Abstract (EN)
Reinforcement learning (RL) is a promising approach to endow robots with skills. However, RL requires many trials to get satisfactory results. Learning from Demonstration (LfD) seeded RL alleviates this problem by learning an initial skill from human demonstrations. Nevertheless, this approach still requires robots to perform a non-trivial amount of trials. In this thesis, we develop an approach to further reduce these for manipulation skills with perceptual goals. Our main contributions are (1) an algorithm to focus the exploration by using the learned relationship between action and perception and a (2) Black-Box RL Policy Search (PS) method that improves upon the popular Policy Improvement with Path Integral (PI²) algorithm, called the Bayesian Optimized PI² (BO-PI²), that uses reward predictive UCB-type exploration. Our underlying LfD framework utilizes a Dynamic Bayesian Network (DBN) learned from keyframe demonstrations to jointly model the action (end-effector pose) and the goal (object-specific perceptual data) of the skill. The action part is used to generate robot trajectories, and the goal part is used to monitor the success of trajectory executions and to create a Partially Observable Markov Reward Model in order to learn rewards. BO-PI$^{2}$ is used to improve the action part of the DBN with trial-and-error using the learned returns. The novelty of BO-PI$^{2}$ comes from its exploration strategy. The coupling between the action and the goal is used to pick the part of the model to focus on to reduce the effort, in a sense to solve the credit attribution problem. After picking the part to focus on, BO-PI$^{2}$ samples trajectories from the action model to get rollouts, which is typical of PS approaches. In addition, BO-PI$^{2}$ uses a Gaussian Process (GP) to learn local returns from these rollouts, which is improved with each executed trajectory. The next samples are selected by utilizing an Upper Confidence Bound (UCB) approach, using the predicted return and uncertainty of the possible candidate points. This is in contrast to random sampling, used in most PS approaches. BO-PI$^{2}$ also utilizes a skill success based termination criteria, using the goal model to monitor success autonomously. We evaluate BO-PI$^{2}$ with expert and non-expert keyframe demonstrations for three skills. In the expert case, the models are perturbed so that the initial skill execution starts from a failure condition. In the non-expert case, we pick skill models that fail to begin with. We test our approach against the current state-of-the-art PI$^{2}$-ES-Cov algorithm using three metrics: (1) skill success rate, (2) total accumulated reward, and (3) number of trials. In both the expert case and the non-expert case, on average, our approach performed better than the baseline on all three metrics. Our results show that utilization of keyframes allows us to focus on failed sub-goals rather than the entire trajectory, and combined with reward predictive exploration strategies, are beneficial to improve RL performance and reduce the number of trials to endow robot arms with real-life manipulation skills.
Author
Dr. Onur Berk Töre
Institution
How to Cite
Onur Berk Töre (Master Thesis). Keyframe demonstration seeded and Bayesian optimized policy search, 2022, Koç University.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Koç University
- Obje tabanlı akıl danışma-tavsiye iletişimi tasarımına ilham kaynağı olarak Türk kahve falı(2017)
- Ekom-Eczacıbaşı'nın Rusya piyasasındaki pazarlama stratejileri(1995)
- Barok döneminde Balkanlar Osmanlı Avrupası'nda mimaride, dekorasyonda, himaye ve kültürel üretim modellerinde dönüşüm, 1718-1856(2006)
- De Rham-Witt kompleks(2011)
- Erteleme kısıtlı tek makine çizelgeleme(2014)
- Sarayda Osmanlı tütsüleme gelenekleri: Topkapı Sarayı buhurdanları(2015)
