Skip to content

Opening book details…

Can I read Minimally Supervised Written-to-Spoken Text Normalization on EtoBox?

Minimally Supervised Written-to-Spoken Text Normalization by Wu, Ke; Gorman, Kyle; Sproat, Richard is a scholarly article available to read on EtoBox.

What is Minimally Supervised Written-to-Spoken Text Normalization about?

In speech-applications such as text-to-speech (TTS) or automatic speech recognition (ASR), \emph{text normalization} refers to the task of converting from a \emph{written} representation into a representation of how the text is to be \emph{spoken}. In all real-world speech applications, the text normalization engine is developed---in large part---by hand. For example, a hand-built grammar may be used to enumerate the possible ways of saying a given token in a given language, and a statistical model used to select the most appropriate pronunciation in context. In this study we examine the tradeoffs associated with using more or less language-specific domain knowledge in a text normalization engine. In the most data-rich scenario, we have access to a carefully constructed hand-built normalization grammar that for any given token will produce a set of all possible verbalizations for that token. We also assume a corpus of aligned written-spoken utterances, from which we can train a ranking model that selects the appropriate verbalization for the given context. As a substitute for the carefully constructed grammar, we also consider a scenario with a language-universal normalization \emp

Author
Wu, Ke; Gorman, Kyle; Sproat, Richard
Published
2016
Language
EN