Moyan AI Training Institution LogoMoyan AI

Research direction · Emerging · Advanced

Vision-language-action models

Also known as: VLA

Models that map camera input and a natural-language instruction directly to robot actions.

What Vision-language-action models is

VLA models aim to give robots the generality of foundation models, so a single policy can follow instructions across tasks and objects it was not explicitly trained on.

How it works

A multimodal backbone is trained on robot demonstration data paired with instructions, outputting action tokens that a controller executes, with safety limits layered on top.

Why it matters

It is the most active route towards general-purpose robots, and progress here is closely watched as a physical-world capability signal.

Common uses

  • Manipulation research
  • Warehouse and lab automation prototypes

Strengths

  • Generalises across tasks
  • Natural-language instruction

Watch for

  • Robot data is scarce
  • Reliability far below industrial requirements

Continue exploring

More in this collection

Browse all AI Models