Visual Speaker Localization Aided by Acoustic Models

The following paper presents a novel audio-visual approach for unsupervised speaker locationing. Using recordings from a single, low-resolution room overview camera and a single far-field microphone, a state-of-the art audio-only speaker localization system (traditionally called speaker diarization) is extended so that both acoustic and visual models are estimated as part of a joint unsupervised optimization problem. The speaker diarization system first automatically determines the number of speakers and estimates “who spoke when”, then, in a second step, the visual models are used to infer the location of the speakers in the video. The experiments were performed on real-world meetings using 4.5 hours of the publicly available AMI meeting corpus. The proposed system is able to exploit audio-visual integration to not only improve the accuracy of a state-of-the-art (audioonly) speaker diarization, but also adds visual speaker locationing at little incremental engineering and computation costs.

Chattez avec Graph Search

Posez n’importe quelle question sur les cours, conférences, exercices, recherches, actualités, etc. de l’EPFL ou essayez les exemples de questions ci-dessous.

AVERTISSEMENT : Le chatbot Graph n'est pas programmé pour fournir des réponses explicites ou catégoriques à vos questions. Il transforme plutôt vos questions en demandes API qui sont distribuées aux différents services informatiques officiellement administrés par l'EPFL. Son but est uniquement de collecter et de recommander des références pertinentes à des contenus que vous pouvez explorer pour vous aider à répondre à vos questions.

Visual Speaker Localization Aided by Acoustic Models

Graph Chatbot

Chattez avec Graph Search

Sound Field Reconstruction in a room through Sparse Recovery and its application in Room Modal Equalization

Learning stereo reconstruction with deep neural networks

Learning and leveraging shared domain semantics to counteract visual domain shifts

Learning and leveraging shared domain semantics to counteract visual domain shifts

Learning stereo reconstruction with deep neural networks

Sound Field Reconstruction in a room through Sparse Recovery and its application in Room Modal Equalization