Publication

Enhancing posterior based speech recognition systems

Hamed Ketabdar
2008
EPFL thesis

Abstract

The use of local phoneme posterior probabilities has been increasingly explored for improving speech recognition systems. Hybrid hidden Markov model / artificial neural network (HMM/ANN) and Tandem are the most successful examples of such systems. In this thesis, we present a principled framework for enhancing the estimation of local posteriors, by integrating phonetic and lexical knowledge, as well as long contextual information. This framework allows for hierarchical estimation, integration and use of local posteriors from the phoneme up to the word level. We propose two approaches for enhancing the posteriors. In the first approach, phoneme posteriors estimated with an ANN (particularly multi-layer Perceptron – MLP) are used as emission probabilities in HMM forward-backward recursions. This yields new enhanced posterior estimates integrating HMM topological constraints (encoding specific phonetic and lexical knowledge), and long context. In the second approach, a temporal context of the regular MLP posteriors is post-processed by a secondary MLP, in order to learn inter and intra dependencies among the phoneme posteriors. The learned knowledge is integrated in the posterior estimation during the inference (forward pass) of the second MLP, resulting in enhanced posteriors. The use of resulting local enhanced posteriors is investigated in a wide range of posterior based speech recognition systems (e.g. Tandem and hybrid HMM/ANN), as a replacement or in combination with the regular MLP posteriors. The enhanced posteriors consistently outperform the regular posteriors in different applications over small and large vocabulary databases.

Official source

https://infoscience.epfl.ch/record/126276?ln=en

About this result

This page is automatically generated and may contain information that is not correct, complete, up-to-date, or relevant to your search query. The same applies to every other page on this website. Please make sure to verify the information with EPFL's official sources.

Graph Chatbot

Chat with Graph Search

Ask any question about EPFL courses, lectures, exercises, research, news, etc. or try the example questions below.

DISCLAIMER: The Graph Chatbot is not programmed to provide explicit or categorical answers to your questions. Rather, it transforms your questions into API requests that are distributed across the various IT services officially administered by EPFL. Its purpose is solely to collect and recommend relevant references to content that you can explore to help you answer your questions.

Hamed Ketabdar
2008
EPFL thesis

Abstract

Official source

https://infoscience.epfl.ch/record/126276?ln=en

About this result

Related concepts (43)

Related publications (70)

Enhancing posterior based speech recognition systems

Graph Chatbot

Chat with Graph Search

Novel Methods For Detection And Analysis Of Atypical Aspects In Speech

On matching data and model in LF-MMI-based dysarthric speech recognition

Novel Methods for Incorporating Prior Knowledge for Automatic Speech Assessment

Novel Methods For Detection And Analysis Of Atypical Aspects In Speech

Novel Methods for Incorporating Prior Knowledge for Automatic Speech Assessment

On matching data and model in LF-MMI-based dysarthric speech recognition