Skip to content

Opening book details…

Can I read UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models on EtoBox?

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models by Li, Yujie; Xu, Wenjia; Li, Guangzuo; Yu, Zijian; Wei, Zhiwei; Wang, Jiuniu; Peng, Mugen is a scholarly article available to read on EtoBox.

What is UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models about?

The domain gap between remote sensing imagery and natural images has recently received widespread attention and Vision-Language Models (VLMs) have demonstrated excellent generalization performance in remote sensing multimodal tasks. However, current research is still limited in exploring how remote sensing VLMs handle different types of visual inputs. To bridge this gap, we introduce \textbf{UniRS}, the first vision-language model \textbf{uni}fying multi-temporal \textbf{r}emote \textbf{s}ensing tasks across various types of visual input. UniRS supports single images, dual-time image pairs, and videos as input, enabling comprehensive remote sensing temporal analysis within a unified framework. We adopt a unified visual representation approach, enabling the model to accept various visual inputs. For dual-time image pair tasks, we customize a change extraction module to further enhance the extraction of spatiotemporal features. Additionally, we design a prompt augmentation mechanism tailored to the model's reasoning process, utilizing the prior knowledge of the general-purpose VLM to provide clues for UniRS. To promote multi-task knowledge sharing, the model is jointly fine-tuned on

Author
Li, Yujie; Xu, Wenjia; Li, Guangzuo; Yu, Zijian; Wei, Zhiwei; Wang, Jiuniu; Peng, Mugen
Published
2024
Language
EN

More by Li, Yujie; Xu, Wenjia; Li, Guangzuo; Yu, Zijian; Wei, Zhiwei; Wang, Jiuniu; Peng, Mugen

Browse all works by Li, Yujie; Xu, Wenjia; Li, Guangzuo; Yu, Zijian; Wei, Zhiwei; Wang, Jiuniu; Peng, Mugen