Sound Event Bounding Boxes

Janek Ebbers; François G. Germain; Gordon Wichern; Jonathan Le Roux

2024 INTERSPEECH INTERSPEECH 2024

Sound Event Bounding Boxes

Abstract

Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence posteriors in short time frames. Then, thresholding produces binary frame-level presence decisions, with the extent of individual events determined by merging presence in consecutive frames. In this paper, we show that frame-level thresholding deteriorates event extent prediction by coupling it with the system’s sound presence confidence. We propose to decouple the prediction of event extent and confidence by introducing sound event bounding boxes (SEBBs), which format each sound event prediction as a combination of a class type, extent, and overall confidence. We also propose a change-detection-based algorithm to convert frame-level posteriors into SEBBs. We find the algorithm significantly improves the performance of DCASE 2023 Challenge systems, boosting the state of the art from .644 to .686 PSDS1.

🧭 Keyword Pioneer — frame-level prediction

🐝 Cross-Pollinator — Artificial Intelligence, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Machine Learning, Mathematics & Optimization, Natural Language Processing, Speech & Audio

🌉 Interdisciplinary Bridge — Artificial Intelligence and Computer Vision and Machine Learning and Speech & Audio

🐣 Hot Topic Early Bird — change detection

Authors

Janek Ebbers , François G. Germain , Gordon Wichern , Jonathan Le Roux

Topics

Machine Learning > Core Methods > Classification Machine Learning > Learning Types > Self-Supervised Learning Computer Vision > Analysis > Object Detection Speech & Audio > Analysis > Speech Analysis Artificial Intelligence > Core AI > Computer Vision

Keywords

change detection posterior estimation bounding box confidence estimation sound event detection frame-level prediction frame-level posterior event extent

Download PDF

Related papers

Reshape Dimensions Network for Speaker Recognition 2024

RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification 2024

Mixed Children/Adult/Childrenized Fine-Tuning for Children’s ASR: How to Reduce Age Mismatch and Speaking Style Mismatch 2024

Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions 2024

K-means and hierarchical clustering of f0 contours 2024