Skip to content

Opening book details…

About this document

End-to-End TTS with Joint Training by tstshin1 is a document available to read on EtoBox.

The document presents a novel end-to-end text-to-speech (E2E-TTS) model that jointly trains FastSpeech2 and HiFi-GAN, simplifying the training pipeline and eliminating the need for external speech-text alignment tools. This model outperforms existing state-of-the-art implementations by synthesizing high-quality speech directly from text without intermediate mel-spectrograms. Key contributions include the integration of an alignment learning objective and the ability to generate speech without requiring pre-

Author
tstshin1
Language
EN