Training & Alignment
24 min
Vision-Language-Action Models: The Action Interface Is the Hard Part
A language model eats trillions of tokens scraped for free. The largest open robot dataset is 527 skills gathered by hand across 21 institutions. That asymmetry, not model capacity, is what makes robot learning hard, and it expla…