ANTI-AI ARCHIVE
ART-HISTORY NODE / 291 · 2024-10-09

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Yu, Sihyun; Kwak, Sangkyung; Jang, Huiwon; Jeong, Jongheon; Huang, Jonathan; Shin, Jinwoo; Xie, Saining

Introduction

REPA aligns intermediate denoising representations with pretrained visual encoders, investigating how visual representations can support diffusion and flow-transformer training.

ORIGINAL DOCUMENT#291
Figure 1: REPA representation alignment and the authors’ training comparison.View full image ↗

Figure 1: REPA representation alignment and the authors’ training comparison. Source: v4; the image version is distinct from the initial submission date.

Yu, Sihyun; Kwak, Sangkyung; Jang, Huiwon; Jeong, Jongheon; Huang, Jonathan; Shin, Jinwoo; Xie, Saining · original paper / arXiv · Rights in the paper and depicted works remain with their holders. Research quotation does not establish an open licence; republication rights await independent review.

Source · Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think ↗

RESEARCH ACCOUNT

REPA aligns intermediate denoising representations with pretrained visual encoders, investigating how visual representations can support diffusion and flow-transformer training.

ANTI-AI ARCHIVE · Revised 2026-10-03

What the paper investigates

A regularisation term aligns projected hidden states for noisy inputs with external encoder representations of clean images. Experiments on architectures including DiT and SiT compare training efficiency and generation quality under this additional supervision.

Section sources: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Its place in generative-art history

Editorial interpretation: visual understanding and image synthesis are not isolated technical traditions. This connection makes the borrowed encoder and its learned representations relevant to a history that might otherwise examine only the final generator.

Section sources: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reading limits and versions

Efficiency and quality claims depend on experimental settings; representation alignment does not establish understanding of cultural meaning.

Section sources: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Sources for this account

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think ↗

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. The account and translation are AI-assisted, pending independent human review. Section references identify evidence without claiming independent verification of every historical statement.

Continue with a comparative question

Scalable Diffusion Models with Transformers ↗

REPA: compare methods, control or evaluation conditions with DiT.

Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis ↗

REPA: compare methods, control or evaluation conditions with Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis ↗

REPA: compare methods, control or evaluation conditions with PixArt-α.

Dates & version record

Date displayed for this node: 2024-10-09 · Historical date recorded for the source: 2024-10-09

The timeline uses the initial arXiv submission, distinct from conference publication, model release and the archive addition on 3 October 2026. The account and image refer to v4; its date is recorded in the source history.

These dates refer to the historical event or recorded version, not this page’s publication date. The original date precision and unresolved questions are retained.

Original sources & further reading

These links lead to the cited paper, article, institution or conference page. External texts retain their source languages.

01
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think ↗

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

https://arxiv.org/abs/2410.06940

Current cited URL

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Provenance, translation & verification

Archive node #291 · Initial research-paper submission; methods, evaluation and production conditions

Research materials & supplement references · 1

Recent research papers: additions and reconciled sources since 2022 · Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. · 027
Date as recorded: 2024-10-09
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think ↗
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Archive account based on the listed research materials, not a full translation of the linked work. AI-assisted translation; independent human review pending.

Official material was read within the stated scope; see the source note for reading limits and outstanding checks.

Cite this node

ANTI-AI ARCHIVE. Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think. Art-history node #291. https://salondesrefuses.cn/en/art-history/350

For specific historical claims, also cite the original sources above and include your access date. This account is not a full translation of the linked work.

Adjacent nodes follow chronological order; adjacency does not establish direct influence or causation.