Pindrop Labs’ Submission to the First Multi-Target Speaker Detection and Identification Challenge

Elie Khoury; Khaled Lakhdhar; Andrew Vaughan; Ganesh Sivaraman; Parav Nagarsheth

2019 INTERSPEECH INTERSPEECH 2019

Pindrop Labs’ Submission to the First Multi-Target Speaker Detection and Identification Challenge

Abstract

This paper summarizes Pindrop Labs’ submission to the multi-target speaker detection and identification challenge evaluation (MCE 2018). The MCE challenge is geared towards detecting blacklisted speakers (fraudsters) in the context of call centers. Particularly, it aims to answer the following two questions: Is the speaker of the test utterance on the blacklist? If so, which speaker is it among the blacklisted speakers? While one single system can answer both questions, this work looks at them as two separate tasks: blacklist detection and closed-set identification. The former is addressed using four different systems including probabilistic linear discriminant analysis (PLDA), two deep neural network (DNN) based systems, and a simple system based on cosine similarity and logistic regression. The latter is addressed by combining PLDA and neural network based systems. The proposed system was the best performing system at the challenge on both tasks, reducing the blacklist detection error (Top-S EER) by 31.9% and the identification error (Top-1 EER) by 46.4% over the MCE baseline on the evaluation data.

🧭 Keyword Pioneer — speaker detection

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Elie Khoury , Khaled Lakhdhar , Andrew Vaughan , Ganesh Sivaraman , Parav Nagarsheth

Topics

Machine Learning > Core Methods > Classification Machine Learning > Core Methods > Metric Learning

Keywords

speaker verification deep neural network speaker identification probabilistic linear discriminant analysis speaker detection blacklist detection

Download PDF

Related papers

Using Real-Time Visual Biofeedback for Second Language Instruction 2019

VAE-Based Regularization for Deep Speaker Embedding 2019

End-to-End SpeakerBeam for Single Channel Target Speech Recognition 2019

Attention-Enhanced Connectionist Temporal Classification for Discrete Speech Emotion Recognition 2019

Attentive to Individual: A Multimodal Emotion Recognition Network with Personalized Attention Profile 2019