Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence Grounding

Jiaming Chen; Weixin Luo; Wei Zhang; Lin Ma

2022 AAAI AAAI 2022

Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence Grounding

Abstract

Abstract Weakly supervised temporal sentence grounding aims to temporally localize the target segment corresponding to a given natural language query, where it provides video-query pairs without temporal annotations during training. Most existing methods use the fused visual-linguistic feature to reconstruct the query, where the least reconstruction error determines the target segment. This work introduces a novel approach that explores the inter-contrast between videos in a composed video by selecting components from two different videos and fusing them into a single video. Such a straightforward yet effective composition strategy provides the temporal annotations at multiple composed positions, resulting in numerous videos with temporal ground-truths for training the temporal sentence grounding task. A transformer framework is introduced with multi-tasks training to learn a compact but efficient visual-linguistic space. The experimental results on the public Charades-STA and ActivityNet-Caption dataset demonstrate the effectiveness of the proposed method, where our approach achieves comparable performance over the state-of-the-art weakly-supervised baselines. The code is available at https://github.com/PPjmchen/Composition_WSTG.

🌉 Interdisciplinary Bridge — Computer Vision and Deep Learning and Machine Learning and Natural Language Processing

🧭 Keyword Pioneer — visual-linguistic space

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Jiaming Chen , Weixin Luo , Wei Zhang , Lin Ma

Topics

Machine Learning > Learning Types > Contrastive Learning Machine Learning > Learning Types > Weakly Supervised Learning Natural Language Processing > Understanding > Semantic Analysis Deep Learning > Learning Types > Multi-Modal Learning Deep Learning > Learning Types > Weakly Supervised Learning Computer Vision > Generation > Visual Question Answering

Keywords

contrastive learning weakly supervised learning video composition temporal sentence grounding visual-linguistic feature visual-linguistic space inter-contrast learning

Download PDF

Related papers

Dynamic Spatial Propagation Network for Depth Completion 2022

FedFR: Joint Optimization Federated Framework for Generic and Personalized Face Recognition 2022

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding 2022

AnchorFace: Boosting TAR@FAR for Practical Face Recognition 2022

Parallel and High-Fidelity Text-to-Lip Generation 2022