2025
ACL
ACL 2025
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset Download PDF
Abstract
AbstractWe introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 91 spoken languages at the intersection of BELEBELE and FLEURS, and one sign language (ASL). As a by-product we also extend the Automatic Speech Recognition Benchmark, FLEURS, by 20%. We evaluate 2M-BELEBELE dataset for both 5-shot and zero-shot settings and across languages, the speech comprehension accuracy is ≈ 10% average lower compared to reading comprehension.
🌉
Interdisciplinary Bridge
— Machine Learning and Speech & Audio
🐝
Cross-Pollinator
— Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Natural Language Processing, Reinforcement Learning, Speech & Audio
🧭
Keyword Pioneer
— sign language understanding
Authors
Topics
Machine Learning > Learning Types > Zero-Shot Learning
Speech & Audio > Recognition > Speech Recognition
Machine Learning > Learning Types > Few-Shot Learning
Speech & Audio > Analysis > Speech Analysis
Deep Learning > Models > Large Language Models
Natural Language Processing > Resources & Methods > Multimodal NLP