Rich Human Feedback for Text-to-Image Generation

Youwei Liang; Junfeng He; Gang Li; Peizhao Li; Arseniy Klimovskiy; Nicholas Carolan; Jiao Sun; Jordi Pont-Tuset; Sarah Young; Feng Yang; Junjie Ke; Krishnamurthy Dj Dvijotham; Katherine M. Collins; Yiwen Luo; Yang Li; Kai J Kohlhoff; Deepak Ramachandran; Vidhya Navalpakkam

2024 CVPR CVPR 2024

Rich Human Feedback for Text-to-Image Generation

Abstract

Recent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However many generated images still suffer from issues such as artifacts/implausibility misalignment with text descriptions and low aesthetic quality. Inspired by the success of Reinforcement Learning with Human Feedback (RLHF) for large language models prior works collected human-provided scores as feedback on generated images and trained a reward model to improve the T2I generation. In this paper we enrich the feedback signal by (i) marking image regions that are implausible or misaligned with the text and (ii) annotating which words in the text prompt are misrepresented or missing on the image. We collect such rich human feedback on 18K generated images (RichHF-18K) and train a multimodal transformer to predict the rich feedback automatically. We show that the predicted rich human feedback can be leveraged to improve image generation for example by selecting high-quality training data to finetune and improve the generative models or by creating masks with predicted heatmaps to inpaint the problematic regions. Notably the improvements generalize to models (Muse) beyond those used to generate the images on which human feedback data were collected (Stable Diffusion variants). The RichHF-18K data set will be released in our GitHub repository: https://github.com/google-research/google-research/tree/master/richhf_18k.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Computer Vision and Deep Learning and Machine Learning and Reinforcement Learning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Youwei Liang , Junfeng He , Gang Li , Peizhao Li , Arseniy Klimovskiy , Nicholas Carolan , Jiao Sun , Jordi Pont-Tuset , Sarah Young , Feng Yang , Junjie Ke , Krishnamurthy Dj Dvijotham , Katherine M. Collins , Yiwen Luo , Yang Li , Kai J Kohlhoff , Deepak Ramachandran , Vidhya Navalpakkam

Topics

Artificial Intelligence > Core AI > Human-AI Interaction Machine Learning > Learning Types > Active Learning Deep Learning > Models > Diffusion Models Computer Vision > Generation > Image Generation Reinforcement Learning > Methods > Deep RL Deep Learning > Learning Types > Reinforcement Learning Deep Learning > Learning Types > Generative Models Deep Learning > Learning Types > Reinforcement Learning from Human Feedback

Keywords

multimodal learning text-to-image generation reinforcement learning from human feedback human feedback reward model multimodal transformer

Download PDF

Related papers

DUSt3R: Geometric 3D Vision Made Easy 2024

Bezier Everywhere All at Once: Learning Drivable Lanes as Bezier Graphs 2024

NeRFDeformer: NeRF Transformation from a Single View via 3D Scene Flows 2024

Unleashing Unlabeled Data: A Paradigm for Cross-View Geo-Localization 2024

DIMAT: Decentralized Iterative Merging-And-Training for Deep Learning Models 2024