Research direction · Emerging · Advanced
Vision-language-action models
Also known as: VLA
Models that map camera input and a natural-language instruction directly to robot actions.
What Vision-language-action models is
VLA models aim to give robots the generality of foundation models, so a single policy can follow instructions across tasks and objects it was not explicitly trained on.
How it works
A multimodal backbone is trained on robot demonstration data paired with instructions, outputting action tokens that a controller executes, with safety limits layered on top.
Why it matters
It is the most active route towards general-purpose robots, and progress here is closely watched as a physical-world capability signal.
Common uses
- →Manipulation research
- →Warehouse and lab automation prototypes
Strengths
- ✓Generalises across tasks
- ✓Natural-language instruction
Watch for
- ✓Robot data is scarce
- ✓Reliability far below industrial requirements
Continue exploring
More in this collection
Browse all AI Models