Opening book details…
Can I read VideoLLM Knows When to Speak: Enhancing Time-Sensitive VideoComprehension with Video-Text Duet Interaction Format on EtoBox?
VideoLLM Knows When to Speak: Enhancing Time-Sensitive VideoComprehension with Video-Text Duet Interaction Format by Yueqian Wang1 Xiaojun Meng2 Yuxuan Wang3Jianxin Lian is a book available to read on EtoBox.
What is VideoLLM Knows When to Speak: Enhancing Time-Sensitive VideoComprehension with Video-Text Duet Interaction Format about?
AbstractRecent researches on video large language models (Vide-oLLM) predominantly focus on model architectures and training datasets, leaving the interaction format between the user and the model under-explored. In existing works, users often interact with VideoLLMs by using the entire video and a query as input, after which the model gener-ates a response. This interaction format constrains the ap-plication of VideoLLMs in scenarios such as live-streaming comprehension where videos do not end and responses are required in a real-time manner, and also results in unsatis-factory performance on time-sensitive tasks that requires lo-calizing video segments. In this paper, we focus on a video-text duet interaction format. This interaction format is char-acterized by the continuous playback of the video, and both the user and the model can insert their text messages at any position during the video playback. When a text message ends, the video continues to play, akin to the alternative of two performers in a duet. We construct MMDuetlT, a video-text training dataset designed to adapt VideoLLMs to video-text duet interaction format. We also introduce the Multi-Answer Grounded Video Oues
- Author
- Yueqian Wang1 Xiaojun Meng2 Yuxuan Wang3Jianxin Lian
- Published
- 2025
- Language
- EN