Skip to content

Opening book details…

About this document

Vision Transformers for Image Recognition by Ali Haider is a document available to read on EtoBox.

The document presents a new method called Vision Transformer (ViT) that applies a standard Transformer architecture directly to sequences of image patches for image recognition tasks. The key findings are: 1) When trained on mid-sized datasets like ImageNet, ViT achieves modest accuracy a few points below comparable ResNet models due to lacking inductive biases like locality and translation equivariance that CNNs provide. 2) However, when pre-trained on larger datasets containing 14M-300M images, the lar

Author
Ali Haider
Language
EN