About this document
9758 Uncertainty Aware Prefere by dunghoang is a document available to read on EtoBox.
The document introduces Diff-UAPA, an algorithm designed to align diffusion policies with human preferences while accounting for uncertainties in preference data. It employs an iterative preference alignment framework and a maximum posterior objective to optimize diffusion policies without requiring explicit reward functions. Extensive experiments demonstrate the robustness of Diff-UAPA in handling diverse and potentially inconsistent human preferences across various decision-making tasks.
- Author
- dunghoang
- Language
- EN