Blind Audio-Visual Source Separation based on Sparse Redundant Representations

Pierre Vandergheynst, Gianluca Monaci, Rémi Gribonval, Anna Llagostera Casanovas
2010
journal articles

Abstract

In this paper we propose a novel method which is able to detect and separate audio-visual sources present in a scene. Our method exploits the correlation between the video signal captured with a camera and a synchronously recorded one-microphone audio track. In a ﬁrst stage, audio and video modalities are decomposed into relevant basic structures using redundant representations. Next, synchrony between relevant events in audio and video modalities is quantiﬁed. Based on this co-occurrence measure, audio-visual sources are counted and located in the image using a robust clustering algorithm that groups video structures exhibiting strong correlations with the audio. Next periods where each source is active alone are determined and used to build spectral Gaussian Mixture Models (GMMs) characterizing the sources acoustic behavior. Finally, these models are used to separate the audio signal in periods during which several sources are mixed. The proposed approach has been extensively tested on synthetic and natural sequences composed of speakers and music instruments. Results show that the proposed method is able to successfully detect, localize, separate and reconstruct present audio-visual sources.

Official source

https://infoscience.epfl.ch/entities/publication/783f0176-53b8-40cd-b556-1bfdc04e14af

About this result

This page is automatically generated and may contain information that is not correct, complete, up-to-date, or relevant to your search query. The same applies to every other page on this website. Please make sure to verify the information with EPFL's official sources.

Blind Audio-Visual Source Separation based on Sparse Redundant Representations

Graph Chatbot

High brightness neutron source optimized for very high pulsed magnetic field experiments

Observation of non-reciprocal harmonic conversion in real sounds

Learnable filter-banks for CNN-based audio applications

High brightness neutron source optimized for very high pulsed magnetic field experiments

Observation of non-reciprocal harmonic conversion in real sounds

Learnable filter-banks for CNN-based audio applications