Can I read Scaling Speech-Text Pre-Training with Synthetic Data on EtoBox?
Scaling Speech-Text Pre-Training with Synthetic Data by hahuyhoanghhh41 is a document available to read on EtoBox.
What is Scaling Speech-Text Pre-Training with Synthetic Data about?
The document presents a novel approach to scaling speech-text pre-training for speech language models (SpeechLMs) using synthetic interleaved data derived from text corpora, which eliminates the need for parallel speech-text datasets. By synthesizing speech tokens from text spans and employing a supervised speech tokenizer, the authors achieved state-of-the-art performance in speech language modeling and developed an end-to-end spoken chatbot. This method significantly enhances the capabilities of SpeechLMs
- Author
- hahuyhoanghhh41
- Language
- EN