Skip to content

Opening book details…

Can I read Crowd-PrefRL: Preference-Based Reward Learning from Crowds on EtoBox?

Crowd-PrefRL: Preference-Based Reward Learning from Crowds by Chhan, David; Novoseller, Ellen; Lawhern, Vernon J. is a scholarly article available to read on EtoBox.

What is Crowd-PrefRL: Preference-Based Reward Learning from Crowds about?

Preference-based reinforcement learning (RL) provides a framework to train AI agents using human feedback through preferences over pairs of behaviors, enabling agents to learn desired behaviors when it is difficult to specify a numerical reward function. While this paradigm leverages human feedback, it typically treats the feedback as given by a single human user. However, different users may desire multiple AI behaviors and modes of interaction. Meanwhile, incorporating preference feedback from crowds (i.e. ensembles of users) in a robust manner remains a challenge, and the problem of training RL agents using feedback from multiple human users remains understudied. In this work, we introduce a conceptual framework, Crowd-PrefRL, that integrates preference-based RL approaches with techniques from unsupervised crowdsourcing to enable training of autonomous system behaviors from crowdsourced feedback. We show preliminary results suggesting that Crowd-PrefRL can learn reward functions and agent policies from preference feedback provided by crowds of unknown expertise and reliability. We also show that in most cases, agents trained with Crowd-PrefRL outperform agents trained with major

Author
Chhan, David; Novoseller, Ellen; Lawhern, Vernon J.
Published
2024
Language
EN

More by Chhan, David; Novoseller, Ellen; Lawhern, Vernon J.

Browse all works by Chhan, David; Novoseller, Ellen; Lawhern, Vernon J.