Master'sOpen Access

Normalizing flows as HMM emissions for learning from demonstration

2022
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Barış Akgün

Abstract (EN)

On top of being used in many different industries, robots are getting out of factories and into our everyday lives in the form of greeter robots, telepresence robots, toys, autonomous cars and perhaps most ubiquitously vacuum cleaners. Soon we may see more capable robots, such as mobile manipulators, helping us in our homes. Programming and controlling robots to achieve certain tasks in controlled industrial environments with field experts is significantly different than using them in everyday environments. Being able to program a robot to achieve desired tasks without the presence of an expert is of importance for the near future. Imitation learning or Learning from Demonstration (LfD) field aims to enable robots to learn from humans. In this framework, instead of analytically deriving and manually programming a skill, the robot learns the desired skill from human demonstrations. Due to the human-interaction aspects, LfD needs to contend with a low amount of demonstrations which leads to a low amount of data. Using reinforcement learning on top of demonstrations is not feasible when there is no one to engineer a reward function, let alone have the setup for trial and error. Thus, extracting as much information as possible from a limited set of demonstrations and utilizing previously learned skills via transfer is an attractive option. In the LfD framework of this thesis, action and goal/perceptual models of skills are learned from keyframes. Action models are used to execute the skill and goal models are used to monitor this execution. Hidden Markov Models (HMMs) and their derivatives are suitable to learn action and goal models from a low amount of keyframe demonstrations. The first part of the thesis introduces the State Traversal Transfer algorithm to facilitate transfer learning of skills for a single user. We show that this algorithm leads to successful transfer compared to learning from scratch for goal models but do not significantly increase action model performance. However, HMMs have some limitations with transfer learning and multiple sources of data (e.g. multiple users, multiple objects for the same skill, etc.), especially about dealing with perceptual states. These limitations partially stem from using multivariate Gaussian emissions and the difficulty of choosing the correct number of hidden states. Towards this end, a generative model called Conditional Flow Hidden Markov Model (C-FlowHMM) is designed by combining conventional HMMs, normalizing flows, and robotic specific adaptations to improve model flexibility in learning goal/perceptual models of skills. The idea is to use a single normalizing flow model, conditioned on hidden states, instead of Gaussians so that a more general emission model can be learned. By using a single model, states share information which is suitable in a low data regime. Furthermore, a neural network model is more amenable to transfer learning. We develop an expectation-maximization (EM) algorithm to train C-FlowHMMs from human demonstrations which lead to better execution monitoring performance compared to HMMs. We also show that C-FlowHMMs result in better transfer learning performance when data is more varied. Finally, we demonstrate that C-FlowHMM is more robust to change in the number of hidden states compared to conventional HMMs.

Author

Farzın Negahbanı

How to Cite

Farzın Negahbanı (Master Thesis). Normalizing flows as HMM emissions for learning from demonstration, 2022, Koç University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Koç University