Are you an EPFL student looking for a semester project?
Work with us on data science and visualisation projects, and deploy your project as an app on top of Graph Search.
In this paper we investigate the detection and recognition of sequences of numbers in spoken utterances. This is done in two steps: first, the entire utterance is decoded assuming that only numbers were spoken. In the second step, non-number segments (garbage) are detected based on word confidence measures. We compare this approach to conventional garbage models. Also, a comparison of several phone posterior based confidence measures is presented in this paper. The work is evaluated in terms of detection task (hit rate and false alarms) and recognition task (word accuracy) within detected number sequences. The proposed method is tested on German continuous spoken utterances where target content (numbers) is only 20%.
Anne-Florence Raphaëlle Bitbol, Damiano Sgarbossa, Umberto Lupo
Grigorios Chrysos, Filippos Kokkinos
Mathieu Salzmann, Frédéric Kaplan, Delphine Ribes Lemay, Nicolas Henchoz, Valentine Bernasconi