Can I read Video Models Are Zero-shot Learners and Reasoners on EtoBox?
Video Models Are Zero-shot Learners and Reasoners by Thaddäus Wiedemer & Yuxuan Li & Paul Vicol & Shixiang Shane Gu & Nick Matarese & Kevin Swersky & Been Kim & Priyank Jaini & Robert Geirhos is a book available to read on EtoBox.
What is Video Models Are Zero-shot Learners and Reasoners about?
The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural languageprocessing from task-specific models to unified, generalist foundation models. This transformationemerged from simple primitives: large, generative models trained on web-scale data. Curiously, thesame primitives apply to today’s generative video models. Could video models be on a trajectorytowards general-purpose vision understanding, much like LLMs developed general-purpose languageunderstanding? We demonstrate that Veo 3 can solve a broad variety of tasks it wasn’t explicitly trainedfor: segmenting objects, detecting edges, editing images, understanding physical properties, recognizingarXiv:2509.20328v1 [cs.LG] 24 Sep 2025object affordances, simulating tool use, and more. These abilities to perceive, model, and manipulathe visual world enable early forms of visual reasoning like maze and symmetry solving. Veo’s emergezero-shot capabilities indicate that video models are on a path to becoming unified, generalist visiofoundation models.Project page: IntroductionWe believe that video models will become unifying, general-purpose foundation models for machinvision just like large langu
- Author
- Thaddäus Wiedemer & Yuxuan Li & Paul Vicol & Shixiang Shane Gu & Nick Matarese & Kevin Swersky & Been Kim & Priyank Jaini & Robert Geirhos
- Published
- 2025
- Language
- EN