About this document
End-to-End TTS with Joint Training by tstshin1 is a document available to read on EtoBox.
The document presents a novel end-to-end text-to-speech (E2E-TTS) model that jointly trains FastSpeech2 and HiFi-GAN, simplifying the training pipeline and eliminating the need for external speech-text alignment tools. This model outperforms existing state-of-the-art implementations by synthesizing high-quality speech directly from text without intermediate mel-spectrograms. Key contributions include the integration of an alignment learning objective and the ability to generate speech without requiring pre-
- Author
- tstshin1
- Language
- EN