Skip to content

Opening book details…

Can I read SALMONN: Towards Generic Hearing Abilities for Large Language Models on EtoBox?

SALMONN: Towards Generic Hearing Abilities for Large Language Models by Tang, Changli; Yu, Wenyi; Sun, Guangzhi; Chen, Xianzhao; Tan, Tian; Li, Wei; Lu, Lu; Ma, Zejun; Zhang, Chao is a scholarly article available to read on EtoBox.

What is SALMONN: Towards Generic Hearing Abilities for Large Language Models about?

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a speech audio language music open neural network, built by integrating a pre-trained text-based large language model (LLM) with speech and audio encoders into a single multimodal model. SALMONN enables the LLM to directly process and understand general audio inputs and achieve competitive performances on a number of speech and audio tasks used in training, such as automatic speech recognition and translation, auditory-information-based question answering, emotion recognition, speaker verification, and music and audio captioning etc. SALMONN also has a diverse set of emergent abilities unseen in the training, which includes but is not limited to speech translation to untrained languages, speech-based slot filling, spoken-query-based question answering, audio-based storytelling, and speech audio co-reasoning etc. The presence of cross-modal emergent abilities is studied, and a novel few-sho

Author
Tang, Changli; Yu, Wenyi; Sun, Guangzhi; Chen, Xianzhao; Tan, Tian; Li, Wei; Lu, Lu; Ma, Zejun; Zhang, Chao
Published
2023
Language
EN