2025
ACL
ACL 2025
NAVER LABS Europe Submission to the Instruction-following Track
Abstract
AbstractIn this paper we describe NAVER LABS Europe submission to the instruction-following speech processing short track at IWSLT 2025. We participate in the constrained settings, developing systems that can simultaneously perform ASR, ST, and SQA tasks from English speech input into the following target languages: Chinese, Italian, and German. Our solution leverages two pretrained modules: (1) a speech-to-LLM embedding projector trained using representations from the SeamlessM4T-v2-large speech encoder; and (2) LoRA adapters trained on text data on top of Llama-3.1-8B-Instruct. These modules are jointly loaded and further instruction-tuned for 1K steps on multilingual and multimodal data to form our final system submitted for evaluation.
🌉
Interdisciplinary Bridge
— Deep Learning and Machine Learning and Speech & Audio
🐝
Cross-Pollinator
— Artificial Intelligence, Computer Science, Computer Vision, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Speech & Audio
Authors
Topics
Artificial Intelligence > Core AI > Multimodal Learning
Machine Learning > Learning Types > Semi-Supervised Learning
Deep Learning > Architectures > Transformers
Natural Language Processing > Resources & Methods > Large Language Models
Speech & Audio > Recognition > Automatic Speech Recognition
Speech & Audio > Recognition > Speech Recognition
Deep Learning > Models > Large Language Models
Artificial Intelligence > Core AI > Multi-Modal Learning