About this document
C V2: Benchmarking Generalization For Vision Language Action Models by shirish.iiser.internship is a document available to read on EtoBox.
C OLOSSEUM V2 is a large-scale simulation benchmark designed to evaluate the generalization capabilities of Vision-Language-Action (VLA) models in robotic manipulation across diverse tasks and conditions. It includes 28 tasks across 13 categories and supports both in-domain and out-of-domain testing, revealing gaps in performance under distribution shifts. The benchmark aims to provide standardized evaluation protocols to facilitate reproducible comparisons and accelerate progress in developing robust robot
- Author
- shirish.iiser.internship
- Language
- EN