Master'sOpen Access

Speech driven backchannel generation in human-robot interaction with conservative Q-learning

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2022
0 views
0 downloads
Advisor: Prof. Dr. Yücel Yemez

Abstract (EN)

Sustaining engagement in human-agent interaction remains an open problem. The purpose of this thesis is to propose a model for maintaining engagement during human-agent interaction through speech-driven backchannel generation. The problem is modeled as a Markov decision process, with the speech signal representing the state and the reward of maximizing human engagement. Due to the fact that online training is frequently impracticable for human-agent interaction, existing datasets on human-to-human dyadic interaction are employed to train an agent for the backchannel generation task. The problem has been addressed using an actor-critic method based on conservative Q-learning (CQL), which reduces the distributional shift problem during training by suppressing Q-value overestimation. The suggested CQL-based approach is objectively evaluated for the laughter generating task on the IEMOCAP dataset. When compared to previous off-policy Q-learning approaches, compliance with the dataset is improved in terms of laugh production rate. Additionally, the learned policy's success is demonstrated by estimating expected engagement with off-policy policy evaluation techniques.

Author

Öykü Zeynep Bayramoğlu

Institution

How to Cite

Öykü Zeynep Bayramoğlu (Master Thesis). Speech driven backchannel generation in human-robot interaction with conservative Q-learning, 2022, Koç University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Koç University