Building Goal-Oriented Dialogue Systems with Situated Visual Context

Sanchit Agarwal; Jan Jezabek; Arijit Biswas; Emre Barut; Bill Gao; Tagyoung Chung

2022 AAAI AAAI 2022

Building Goal-Oriented Dialogue Systems with Situated Visual Context

Abstract

Abstract Goal-oriented dialogue agents can comfortably utilize the conversational context and understand its users' goals. However, in visually driven user experiences, these conversational agents are also required to make sense of the screen context in order to provide a proper interactive experience. In this paper, we propose a novel multimodal conversational framework where the dialogue agent's next action and their arguments are derived jointly conditioned both on the conversational and the visual context. We demonstrate the proposed approach via a prototypical furniture shopping experience for a multimodal virtual assistant.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Computer Vision and Deep Learning and Natural Language Processing

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Sanchit Agarwal , Jan Jezabek , Arijit Biswas , Emre Barut , Bill Gao , Tagyoung Chung

Topics

Computer Vision > Analysis > Scene Understanding Computer Vision > Processing > Video Understanding Natural Language Processing > Generation > Dialogue Systems Natural Language Processing > Applications > Dialogue Systems Deep Learning > Learning Types > Multimodal Learning Artificial Intelligence > Core AI > Dialogue Systems

Keywords

multimodal learning visual context goal-oriented dialogue dialogue system dialogue agent multimodal conversation virtual assistant

Download PDF

Related papers

Dynamic Spatial Propagation Network for Depth Completion 2022

FedFR: Joint Optimization Federated Framework for Generic and Personalized Face Recognition 2022

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding 2022

AnchorFace: Boosting TAR@FAR for Practical Face Recognition 2022

Parallel and High-Fidelity Text-to-Lip Generation 2022