Segmental Encoder-Decoder Models for Large Vocabulary Automatic Speech Recognition

Eugen Beck; Mirko Hannemann; Patrick Dötsch; Ralf Schlüter; Hermann Ney

2018 INTERSPEECH INTERSPEECH 2018

Segmental Encoder-Decoder Models for Large Vocabulary Automatic Speech Recognition

Abstract

It has been known for a long time that the classic Hidden-Markov-Model (HMM) derivation for speech recognition contains assumptions such as independence of observation vectors and weak duration modeling that are practical but unrealistic. When using the hybrid approach this is amplified by trying to fit a discriminative model into a generative one. Hidden Conditional Random Fields (CRFs) and segmental models (e.g. Semi-Markov CRFs / Segmental CRFs) have been proposed as an alternative, but for a long time have failed to get traction until recently. In this paper we explore different length modeling approaches for segmental models, their relation to attention-based systems. Furthermore we show experimental results on a handwriting recognition task and to the best of our knowledge the first reported results on the Switchboard 300h speech recognition corpus using this approach.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Speech & Audio

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Eugen Beck , Mirko Hannemann , Patrick Dötsch , Ralf Schlüter , Hermann Ney

Topics

Artificial Intelligence > Core AI > Foundation Models Speech & Audio > Recognition > Automatic Speech Recognition

Keywords

automatic speech recognition hidden markov model conditional random field encoder-decoder model segmental model large vocabulary

Download PDF

Related papers

HoloCompanion: An MR Friend for EveryOne 2018

Estimation of the Vocal Tract Length of Vowel Sounds Based on the Frequency of the Significant Spectral Valley 2018

Deep Learning Techniques for Koala Activity Detection 2018

An Exploration of Local Speaking Rate Variations in Mandarin Read Speech 2018

Acoustic Analysis of Whispery Voice Disguise in Mandarin Chinese 2018