Opening book details…
Can I read Moshi: a Speech-text Foundation Model for Real-time Dialogue on EtoBox?
Moshi: a Speech-text Foundation Model for Real-time Dialogue by Défossez, Alexandre; Mazaré, Laurent; Orsini, Manu; Royer, Amélie; Pérez, Patrick; Jégou, Hervé; Grave, Edouard; Zeghidour, Neil is a scholarly article available to read on EtoBox.
What is Moshi: a Speech-text Foundation Model for Real-time Dialogue about?
We introduce Moshi, a speech-text foundation model and full-duplex spoken dialogue framework. Current systems for spoken dialogue rely on pipelines of independent components, namely voice activity detection, speech recognition, textual dialogue and text-to-speech. Such frameworks cannot emulate the experience of real conversations. First, their complexity induces a latency of several seconds between interactions. Second, text being the intermediate modality for dialogue, non-linguistic information that modifies meaning -- such as emotion or non-speech sounds -- is lost in the interaction. Finally, they rely on a segmentation into speaker turns, which does not take into account overlapping speech, interruptions and interjections. Moshi solves these independent issues altogether by casting spoken dialogue as speech-to-speech generation. Starting from a text language model backbone, Moshi generates speech as tokens from the residual quantizer of a neural audio codec, while modeling separately its own speech and that of the user into parallel streams. This allows for the removal of explicit speaker turns, and the modeling of arbitrary conversational dynamics. We moreover extend the hie
- Author
- Défossez, Alexandre; Mazaré, Laurent; Orsini, Manu; Royer, Amélie; Pérez, Patrick; Jégou, Hervé; Grave, Edouard; Zeghidour, Neil
- Published
- 2024
- Language
- EN