Skip to content

Opening book details…

Can I read CT-MVSNet: Efficient Multi-View Stereo with Cross-scale Transformer on EtoBox?

CT-MVSNet: Efficient Multi-View Stereo with Cross-scale Transformer by Wang, Sicheng; Jiang, Hao; Xiang, Lei is a scholarly article available to read on EtoBox.

What is CT-MVSNet: Efficient Multi-View Stereo with Cross-scale Transformer about?

Recent deep multi-view stereo (MVS) methods have widely incorporated transformers into cascade network for high-resolution depth estimation, achieving impressive results. However, existing transformer-based methods are constrained by their computational costs, preventing their extension to finer stages. In this paper, we propose a novel cross-scale transformer (CT) that processes feature representations at different stages without additional computation. Specifically, we introduce an adaptive matching-aware transformer (AMT) that employs different interactive attention combinations at multiple scales. This combined strategy enables our network to capture intra-image context information and enhance inter-image feature relationships. Besides, we present a dual-feature guided aggregation (DFGA) that embeds the coarse global semantic information into the finer cost volume construction to further strengthen global and local feature awareness. Meanwhile, we design a feature metric loss (FM Loss) that evaluates the feature bias before and after transformation to reduce the impact of feature mismatch on depth estimation. Extensive experiments on DTU dataset and Tanks and Temples (T\&T) ben

Author
Wang, Sicheng; Jiang, Hao; Xiang, Lei
Published
2023
Language
EN

More by Wang, Sicheng; Jiang, Hao; Xiang, Lei

Browse all works by Wang, Sicheng; Jiang, Hao; Xiang, Lei