Text-Based Interactive Recommendation via Offline Reinforcement Learning

Ruiyi Zhang; Tong Yu; Yilin Shen; Hongxia Jin

2022 AAAI AAAI 2022

Text-Based Interactive Recommendation via Offline Reinforcement Learning

Abstract

Abstract Interactive recommendation with natural-language feedback can provide richer user feedback and has demonstrated advantages over traditional recommender systems. However, the classical online paradigm involves iteratively collecting experience via interaction with users, which is expensive and risky. We consider an offline interactive recommendation to exploit arbitrary experience collected by multiple unknown policies. A direct application of policy learning with such fixed experience suffers from the distribution shift. To tackle this issue, we develop a behavior-agnostic off-policy correction framework to make offline interactive recommendation possible. Specifically, we leverage the conservative Q-function to perform off-policy evaluation, which enables learning effective policies from fixed datasets without further interactions. Empirical results on the simulator derived from real-world datasets demonstrate the effectiveness of our proposed offline training framework.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Data Science & Analytics and Machine Learning and Reinforcement Learning

🧭 Keyword Pioneer — conservative q-function

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Ruiyi Zhang , Tong Yu , Yilin Shen , Hongxia Jin

Topics

Artificial Intelligence > Learning Paradigms > Transfer Learning Reinforcement Learning > Methods > Offline RL Data Science & Analytics > Applications > Recommender Systems Machine Learning > Learning Types > Reinforcement Learning Machine Learning > Learning Types > Offline Reinforcement Learning

Keywords

offline reinforcement learning distribution shift recommender system off-policy correction interactive recommendation natural language feedback conservative q-function

Download PDF

Related papers

Dynamic Spatial Propagation Network for Depth Completion 2022

FedFR: Joint Optimization Federated Framework for Generic and Personalized Face Recognition 2022

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding 2022

AnchorFace: Boosting TAR@FAR for Practical Face Recognition 2022

Parallel and High-Fidelity Text-to-Lip Generation 2022