arxivcs.CVcs.AI2026-06-25
Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
Yunqi Xue, Zhijiang Li, Philip Torr, Jindong Gu
Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens. These tokens are derived from a codebook that maps embeddings to quantized visual patterns. The la…