Skip to content

Opening book details…

Can I read GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement on EtoBox?

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement by Yang, Yifan; Song, Zheshu; Zhuo, Jianheng; Cui, Mingyu; Li, Jinpeng; Yang, Bo; Du, Yexing; Ma, Ziyang; Liu, Xunying; Wang, Ziyuan; Li, Ke; Fan, Shuai; Yu, Kai; Zhang, Wei-Qiang; Chen, Guoguo; Chen, Xie is a scholarly article available to read on EtoBox.

What is GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement about?

The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual speech recognition corpus. It is designed for low-resource languages and does not rely on paired speech and text data. GigaSpeech 2 comprises about 30,000 hours of automatically transcribed speech, including Thai, Indonesian, and Vietnamese, gathered from unlabeled YouTube videos. We also introduce an automated pipeline for data crawling, transcription, and label refinement. Specifically, this pipeline uses Whisper for initial transcription and TorchAudio for forced alignment, combined with multi-dimensional filtering for data quality assurance. A modified Noisy Student Training is developed to further refine flawed pseudo labels iteratively, thus enhancing model performance. Experimental results on our manually transcribed evaluation set and two public test sets from Common Voice and FLEURS confirm our corpus's high quality and broad applicability. Notably, ASR models trained on GigaSpee

Author
Yang, Yifan; Song, Zheshu; Zhuo, Jianheng; Cui, Mingyu; Li, Jinpeng; Yang, Bo; Du, Yexing; Ma, Ziyang; Liu, Xunying; Wang, Ziyuan; Li, Ke; Fan, Shuai; Yu, Kai; Zhang, Wei-Qiang; Chen, Guoguo; Chen, Xie
Published
2024
Language
EN

More by Yang, Yifan; Song, Zheshu; Zhuo, Jianheng; Cui, Mingyu; Li, Jinpeng; Yang, Bo; Du, Yexing; Ma, Ziyang; Liu, Xunying; Wang, Ziyuan; Li, Ke; Fan, Shuai; Yu, Kai; Zhang, Wei-Qiang; Chen, Guoguo; Chen, Xie

Browse all works by Yang, Yifan; Song, Zheshu; Zhuo, Jianheng; Cui, Mingyu; Li, Jinpeng; Yang, Bo; Du, Yexing; Ma, Ziyang; Liu, Xunying; Wang, Ziyuan; Li, Ke; Fan, Shuai; Yu, Kai; Zhang, Wei-Qiang; Chen, Guoguo; Chen, Xie