About this document
CVPR22 Igv by lyc071719 is a document available to read on EtoBox.
This document discusses a new framework called Invariant Grounding for Video Question Answering (IGV), which aims to improve the reasoning ability of VideoQA models by addressing the issue of spurious correlations between visual scenes and answers. The authors argue that traditional methods, which rely on empirical risk minimization, fail to distinguish between causal scenes and irrelevant complements, leading to unreliable answers. Through experiments on benchmark datasets, IGV demonstrates enhanced accura
- Author
- lyc071719
- Language
- EN