Neural VTLN for Speaker Adaptation in TTS

Chat with Graph Search

Ask any question about EPFL courses, lectures, exercises, research, news, etc. or try the example questions below.

DISCLAIMER: The Graph Chatbot is not programmed to provide explicit or categorical answers to your questions. Rather, it transforms your questions into API requests that are distributed across the various IT services officially administered by EPFL. Its purpose is solely to collect and recommend relevant references to content that you can explore to help you answer your questions.

Vocal tract length normalisation (VTLN) is well established as a speaker adaptation technique that can work with very little adaptation data. It is also well known that VTLN can be cast as a linear transform in the cepstral domain. Building on this latter property, we show that it can be cast as a (linear) layer in a deep neural network (DNN) for speech synthesis. We show that VTLN parameters can then be trained in the same framework as the rest of the DNN using automatic gradients. Experimental results show that the DNN is capable of predicting phone-dependent warpings on artificial data, and that such warpings improve the quality of an acoustic model on real data in subjective listening tests.

Neural VTLN for Speaker Adaptation in TTS

Graph Chatbot

Chat with Graph Search

Infusing structured knowledge priors in neural models for sample-efficient symbolic reasoning

Generalization of Scaled Deep ResNets in the Mean-Field Regime

Coupling a recurrent neural network to SPAD TCSPC systems for real-time fluorescence lifetime imaging

Infusing structured knowledge priors in neural models for sample-efficient symbolic reasoning

Generalization of Scaled Deep ResNets in the Mean-Field Regime

Coupling a recurrent neural network to SPAD TCSPC systems for real-time fluorescence lifetime imaging