Diffusion Model Alignment Using Direct Preference Optimization
Published • Jun 16, 2024
Authors:,,
Bram Wallace
Meihua Dang
Rafael Rafailov
Abstract
Large language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' preferences. In contrast to LLMs, human preference learning has not been widely explored in text-to-image diffusion models; t...
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!