Skip to content

Opening book details…

Can I read VideoLLM Knows When to Speak: Enhancing Time-Sensitive VideoComprehension with Video-Text Duet Interaction Format on EtoBox?

VideoLLM Knows When to Speak: Enhancing Time-Sensitive VideoComprehension with Video-Text Duet Interaction Format by Yueqian Wang1 Xiaojun Meng2 Yuxuan Wang3Jianxin Lian is a book available to read on EtoBox.

What is VideoLLM Knows When to Speak: Enhancing Time-Sensitive VideoComprehension with Video-Text Duet Interaction Format about?

AbstractRecent researches on video large language models (Vide-oLLM) predominantly focus on model architectures and training datasets, leaving the interaction format between the user and the model under-explored. In existing works, users often interact with VideoLLMs by using the entire video and a query as input, after which the model gener-ates a response. This interaction format constrains the ap-plication of VideoLLMs in scenarios such as live-streaming comprehension where videos do not end and responses are required in a real-time manner, and also results in unsatis-factory performance on time-sensitive tasks that requires lo-calizing video segments. In this paper, we focus on a video-text duet interaction format. This interaction format is char-acterized by the continuous playback of the video, and both the user and the model can insert their text messages at any position during the video playback. When a text message ends, the video continues to play, akin to the alternative of two performers in a duet. We construct MMDuetlT, a video-text training dataset designed to adapt VideoLLMs to video-text duet interaction format. We also introduce the Multi-Answer Grounded Video Oues

Author
Yueqian Wang1 Xiaojun Meng2 Yuxuan Wang3Jianxin Lian
Published
2025
Language
EN