About this document
Vision Transformers for Image Recognition by Ali Haider is a document available to read on EtoBox.
The document presents a new method called Vision Transformer (ViT) that applies a standard Transformer architecture directly to sequences of image patches for image recognition tasks. The key findings are: 1) When trained on mid-sized datasets like ImageNet, ViT achieves modest accuracy a few points below comparable ResNet models due to lacking inductive biases like locality and translation equivariance that CNNs provide. 2) However, when pre-trained on larger datasets containing 14M-300M images, the lar
- Author
- Ali Haider
- Language
- EN