ANTI-AI ARCHIVE
ART-HISTORY NODE / 268 · 2024-04-03

Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Tian, Keyu; Jiang, Yi; Yuan, Zehuan; Peng, Bingyue; Wang, Liwei

Introduction

VAR reformulates visual autoregression as coarse-to-fine next-scale prediction, providing a distinct image-generation route to compare with diffusion transformers.

ORIGINAL DOCUMENT#268
Figure 1: VAR generation and zero-shot editing examples.View full image ↗

Figure 1: VAR generation and zero-shot editing examples. Source: v2; the image version is distinct from the initial submission date.

Tian, Keyu; Jiang, Yi; Yuan, Zehuan; Peng, Bingyue; Wang, Liwei · original paper / arXiv · Rights in the paper and depicted works remain with their holders. Research quotation does not establish an open licence; republication rights await independent review.

Source · Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction ↗

RESEARCH ACCOUNT

VAR reformulates visual autoregression as coarse-to-fine next-scale prediction, providing a distinct image-generation route to compare with diffusion transformers.

ANTI-AI ARCHIVE · Revised 2026-10-03

What the paper investigates

The representation is organised by resolution, with each stage predicting a finer scale. The paper studies scaling, ImageNet generation and transfer to tasks including inpainting and outpainting, and releases models and code for investigation.

Section sources: Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Its place in generative-art history

Editorial interpretation: coarse-to-fine formation offers a different process from raster scanning or gradual denoising. It supports comparison of how machines organise image structure without treating language-like token order as the only autoregressive form.

Section sources: Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reading limits and versions

Performance comparisons concern specified models and benchmarks, not universal superiority of autoregression over diffusion.

Section sources: Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Sources for this account

Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction ↗

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. The account and translation are AI-assisted, pending independent human review. Section references identify evidence without claiming independent verification of every historical statement.

Continue with a comparative question

Autoregressive Image Generation without Vector Quantization ↗

VAR: compare methods, control or evaluation conditions with MAR / Diffusion Loss.

PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis ↗

VAR: compare methods, control or evaluation conditions with PixArt-α.

Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis ↗

VAR: compare methods, control or evaluation conditions with Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Dates & version record

Date displayed for this node: 2024-04-03 · Historical date recorded for the source: 2024-04-03

The timeline uses the initial arXiv submission, distinct from conference publication, model release and the archive addition on 3 October 2026. The account and image refer to v2; its date is recorded in the source history.

These dates refer to the historical event or recorded version, not this page’s publication date. The original date precision and unresolved questions are retained.

Original sources & further reading

These links lead to the cited paper, article, institution or conference page. External texts retain their source languages.

01
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction ↗

Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

https://arxiv.org/abs/2404.02905

Current cited URL

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Provenance, translation & verification

Archive node #268 · Initial research-paper submission; methods, evaluation and production conditions

Research materials & supplement references · 1

Recent research papers: additions and reconciled sources since 2022 · Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. · 025
Date as recorded: 2024-04-03
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction ↗
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Archive account based on the listed research materials, not a full translation of the linked work. AI-assisted translation; independent human review pending.

Official material was read within the stated scope; see the source note for reading limits and outstanding checks.

Cite this node

ANTI-AI ARCHIVE. Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction. Art-history node #268. https://salondesrefuses.cn/en/art-history/348

For specific historical claims, also cite the original sources above and include your access date. This account is not a full translation of the linked work.

Adjacent nodes follow chronological order; adjacency does not establish direct influence or causation.