A Bayesian Approach To Inter-Task Fusion For Speaker Recognition

In i-vector based speaker recognition systems, back-end classifiers are trained to factor out nuisance information and retain only the speaker identity. As a result, variabilities arising due to gender, language and accent ( among many others) are suppressed. Inter-task fusion, in which such metadata information obtained from automatic systems is used, has been shown to improve speaker recognition performance. In this paper, we explore a Bayesian approach towards inter-task fusion. Speaker similarity score for a test recording is obtained by marginalizing the posterior probability of a speaker. Gender and language probabilities for the test audio are combined with speaker posteriors to obtain a final speaker score. The proposed approach is demonstrated for speaker verification and speaker identification tasks on the NIST SRE 2008 dataset. Relative improvements of up to 10% and 8% are obtained when fusing gender and language information, respectively.

Chattez avec Graph Search

Posez n’importe quelle question sur les cours, conférences, exercices, recherches, actualités, etc. de l’EPFL ou essayez les exemples de questions ci-dessous.

AVERTISSEMENT : Le chatbot Graph n'est pas programmé pour fournir des réponses explicites ou catégoriques à vos questions. Il transforme plutôt vos questions en demandes API qui sont distribuées aux différents services informatiques officiellement administrés par l'EPFL. Son but est uniquement de collecter et de recommander des références pertinentes à des contenus que vous pouvez explorer pour vous aider à répondre à vos questions.

A Bayesian Approach To Inter-Task Fusion For Speaker Recognition

Graph Chatbot

Chattez avec Graph Search

A BAYESIAN APPROACH TO INTER-TASK FUSION FOR SPEAKER RECOGNITION

Efficient learning of smooth probability functions from Bernoulli tests with guarantees.

Comparing methodologies for structural identification and fatigue life prediction of a highway bridge

A BAYESIAN APPROACH TO INTER-TASK FUSION FOR SPEAKER RECOGNITION

Efficient learning of smooth probability functions from Bernoulli tests with guarantees.

Comparing methodologies for structural identification and fatigue life prediction of a highway bridge