End-to-End Joint Target and Non-Target Speakers ASR

Ryo Masumura; Naoki Makishima; Taiga Yamane; Yoshihiko Yamazaki; Saki Mizuno; Mana Ihori; Mihiro Uchida; Keita Suzuki; Hiroshi Sato; Tomohiro Tanaka; Akihiko Takashima; Satoshi Suzuki; Takafumi Moriya; Nobukatsu Hojo; Atsushi Ando

2023 INTERSPEECH INTERSPEECH 2023

End-to-End Joint Target and Non-Target Speakers ASR

Abstract

This paper proposes a novel automatic speech recognition (ASR) system that can transcribe individual speaker's speech while identifying whether they are target or non-target speakers from multi-talker overlapped speech. Target-speaker ASR systems are a promising way to only transcribe a target speaker's speech by enrolling the target speaker's information. However, in conversational ASR applications, transcribing both the target speaker's speech and non-target speakers' ones is often required to understand interactive information. To naturally consider both target and non-target speakers in a single ASR model, our idea is to extend autoregressive modeling-based multi-talker ASR systems to utilize the enrollment speech of the target speaker. Our proposed ASR is performed by recursively generating both textual tokens and tokens that represent target or non-target speakers. Our experiments demonstrate the effectiveness of our proposed method.

🌉 Interdisciplinary Bridge — Deep Learning and Speech & Audio

🧭 Keyword Pioneer — target speaker identification

🐝 Cross-Pollinator — Artificial Intelligence, Deep Learning, Machine Learning, Natural Language Processing, Speech & Audio

Authors

Ryo Masumura , Naoki Makishima , Taiga Yamane , Yoshihiko Yamazaki , Saki Mizuno , Mana Ihori , Mihiro Uchida , Keita Suzuki , Hiroshi Sato , Tomohiro Tanaka , Akihiko Takashima , Satoshi Suzuki , Takafumi Moriya , Nobukatsu Hojo , Atsushi Ando

Topics

Deep Learning > Architectures > Neural Networks Speech & Audio Speech & Audio > Recognition > Automatic Speech Recognition

Keywords

automatic speech recognition multi-talker speech recognition end-to-end asr target speaker identification overlapped speech speaker enrollment autoregressive modeling multi-talker speech target speaker

Download PDF

Audio-Visual Praise Estimation for Conversational Video based on Synchronization-Guided Multimodal Transformer 2023

Improving the response timing estimation for spoken dialogue systems by reducing the effect of speech recognition delay 2023

Improving Code-Switching and Name Entity Recognition in ASR with Speech Editing based Data Augmentation 2023

What are differences? Comparing DNN and Human by Their Performance and Characteristics in Speaker Age Estimation 2023

End-to-End Joint Target and Non-Target Speakers ASR

Abstract

Authors

Topics

Keywords

Related papers