Can I read Huang AV-Mamba Cross-Modality Selective State Space Models For Audio-Visual Question Answering on EtoBox?
Huang AV-Mamba Cross-Modality Selective State Space Models For Audio-Visual Question Answering by shwetas is a document available to read on EtoBox.
What is Huang AV-Mamba Cross-Modality Selective State Space Models For Audio-Visual Question Answering about?
The document presents AV-Mamba, a novel framework for audio-visual question answering (AVQA) that enhances the selection of relevant information across audio, visual, and textual modalities using a cross-modality selection mechanism. The framework incorporates components for feature extraction, audio-guided spatial grounding, and question-guided temporal grounding, demonstrating improved performance on the MUSIC-AVQA dataset compared to existing models. Extensive experiments validate the effectiveness of AV
- Author
- shwetas
- Language
- EN