NobleBlocks
Public

Diffusion Model Alignment Using Direct Preference Optimization

Published • Jun 16, 2024
Authors:
Bram Wallace
,
Meihua Dang
,
Rafael Rafailov

Abstract

Large language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' preferences. In contrast to LLMs, human preference learning has not been widely explored in text-to-image diffusion models; t...

Finding related papers...

Discussions

(0)

No comments yet

Be the first to share your thoughts!