ANTI-AI ARCHIVE
ART-HISTORY NODE / 207 · 2022-12-19

Scalable Diffusion Models with Transformers

Scalable Diffusion Models with Transformers

Peebles, William; Xie, Saining

Introduction

DiT replaces the usual U-Net backbone with a transformer operating on latent patches and studies compute scaling, providing a key architectural comparison for later generative models.

ORIGINAL DOCUMENT#207
Figure 1: samples from class-conditioned DiT models.View full image ↗

Figure 1: samples from class-conditioned DiT models. Source: v2; the image version is distinct from the initial submission date.

Peebles, William; Xie, Saining · original paper / arXiv · Rights in the paper and depicted works remain with their holders. Research quotation does not establish an open licence; republication rights await independent review.

Source · Scalable Diffusion Models with Transformers ↗

RESEARCH ACCOUNT

DiT replaces the usual U-Net backbone with a transformer operating on latent patches and studies compute scaling, providing a key architectural comparison for later generative models.

ANTI-AI ARCHIVE · Revised 2026-10-03

What the paper investigates

Peebles and Xie replace the diffusion backbone with a transformer processing patches of image latents. Comparisons vary depth, width and token count. Class-conditioned ImageNet experiments relate forward-pass computation to FID, rather than testing a complete consumer text-to-image application.

Section sources: Scalable Diffusion Models with Transformers

Its place in generative-art history

Editorial interpretation: this fills an architectural link between diffusion methods and scalable visual models. Reading it with PixArt and SD3 allows language conditioning, network structure and training scale to be tracked separately instead of compressing all changes into a product name.

Section sources: Scalable Diffusion Models with Transformers

Reading limits and versions

2022 is the preprint year. Class-conditioned benchmarks do not directly establish text understanding or the artistic value of outputs.

Section sources: Scalable Diffusion Models with Transformers

Sources for this account

Scalable Diffusion Models with Transformers ↗

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. The account and translation are AI-assisted, pending independent human review. Section references identify evidence without claiming independent verification of every historical statement.

Continue with a comparative question

Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction ↗

DiT: compare methods, control or evaluation conditions with VAR.

PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis ↗

DiT: compare methods, control or evaluation conditions with PixArt-α.

Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis ↗

DiT: compare methods, control or evaluation conditions with Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Dates & version record

Date displayed for this node: 2022-12-19 · Historical date recorded for the source: 2022-12-19

The timeline uses the initial arXiv submission, distinct from conference publication, model release and the archive addition on 3 October 2026. The account and image refer to v2; its date is recorded in the source history.

These dates refer to the historical event or recorded version, not this page’s publication date. The original date precision and unresolved questions are retained.

Original sources & further reading

These links lead to the cited paper, article, institution or conference page. External texts retain their source languages.

01
Scalable Diffusion Models with Transformers ↗

Scalable Diffusion Models with Transformers

https://arxiv.org/abs/2212.09748

Current cited URL

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Provenance, translation & verification

Archive node #207 · Initial research-paper submission; methods, evaluation and production conditions

Research materials & supplement references · 1

Recent research papers: additions and reconciled sources since 2022 · Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. · 009
Date as recorded: 2022-12-19
Scalable Diffusion Models with Transformers ↗
Scalable Diffusion Models with Transformers

Archive account based on the listed research materials, not a full translation of the linked work. AI-assisted translation; independent human review pending.

Official material was read within the stated scope; see the source note for reading limits and outstanding checks.

Cite this node

ANTI-AI ARCHIVE. Scalable Diffusion Models with Transformers. Art-history node #207. https://salondesrefuses.cn/en/art-history/332

For specific historical claims, also cite the original sources above and include your access date. This account is not a full translation of the linked work.

Adjacent nodes follow chronological order; adjacency does not establish direct influence or causation.