Multimodal AI · Emerging · Advanced
Continuous Autoregressive Vision
Also known as: Autoregressive Image Models, Visual Autoregression
An architecture that generates images token-by-token or scale-by-scale using autoregressive transformer mechanics.
What Continuous Autoregressive Vision is
Provides unified multimodal architectures where text and image generation share identical transformer underlying logic.
How it works
Quantizes visual tokens using vector-quantized autoencoders and predicts next visual tokens via causal attention.
Why it matters
Pioneered in models like LLaMA-Vision, Chameleon, and Emu3 for seamless multimodal synthesis.
Common uses
- →Unified text-image foundation models
- →Visual content synthesis
- →Interleaved multimodal generation
More in this collection
Browse all AI Concepts