Skip to content

Opening book details…

Can I read Huang AV-Mamba Cross-Modality Selective State Space Models For Audio-Visual Question Answering on EtoBox?

Huang AV-Mamba Cross-Modality Selective State Space Models For Audio-Visual Question Answering by shwetas is a document available to read on EtoBox.

What is Huang AV-Mamba Cross-Modality Selective State Space Models For Audio-Visual Question Answering about?

The document presents AV-Mamba, a novel framework for audio-visual question answering (AVQA) that enhances the selection of relevant information across audio, visual, and textual modalities using a cross-modality selection mechanism. The framework incorporates components for feature extraction, audio-guided spatial grounding, and question-guided temporal grounding, demonstrating improved performance on the MUSIC-AVQA dataset compared to existing models. Extensive experiments validate the effectiveness of AV

Author
shwetas
Language
EN