Skip to content

Opening book details…

Can I read Bi-Modal Transformer-Based Approach For Visual Question Answering in Remote Sensing Imagery on EtoBox?

Bi-Modal Transformer-Based Approach For Visual Question Answering in Remote Sensing Imagery by Ayushmaan Lohani is a document available to read on EtoBox.

What is Bi-Modal Transformer-Based Approach For Visual Question Answering in Remote Sensing Imagery about?

This document presents a novel visual question answering (VQA) approach for remote sensing imagery using bi-modal transformer-based models. The proposed method leverages the contrastive language image pretraining (CLIP) network to embed visual and textual representations, and employs attention mechanisms to capture their interdependencies. Experimental results demonstrate that this approach achieves competitive performance with reduced training data on datasets from Sentinel-2 and aerial sensors.

Author
Ayushmaan Lohani
Language
EN