Skip to content

Opening book details…

Can I read Speaker and Expression Factorization for Audiobook Data: Expressiveness and Transplantation on EtoBox?

Speaker and Expression Factorization for Audiobook Data: Expressiveness and Transplantation by Langzhou Chen; Norbert Braunschweiler; Mark J. F. Gales is a scholarly article available to read on EtoBox.

What is Speaker and Expression Factorization for Audiobook Data: Expressiveness and Transplantation about?

Expressive synthesis from text is a challenging problem. There are two issues. First, read text is often highly expressive to convey the emotion and scenario in the text. Second, since the expressive training speech is not always available for different speakers, it is necessary to develop methods to share the expressive information over speakers. This paper investigates the approach of using very expressive, highly diverse audiobook data from multiple speakers to build an expressive speech synthesis system. Both of two problems are addressed by considering a factorized framework where speaker and emotion are modeled in separate sub-spaces of a cluster adaptive training (CAT) parametric speech synthesis system. The sub-spaces for the expressive state of a speaker and the characteristics of the speaker are jointly trained using a set of audiobooks. In this work, the expressive speech synthesis system works in two distinct modes. In the first mode, the expressive information is given by audio data and the adaptation method is used to extract the expressive information in the audio data. In the second mode, the input of the synthesis system is plain text and a full expressive synthesi

Author
Langzhou Chen; Norbert Braunschweiler; Mark J. F. Gales
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Published
2015
Language
EN