ANTI-AI ARCHIVE
ART-HISTORY NODE / 210 · 2023-01-02

Muse: Text-To-Image Generation via Masked Generative Transformers

Muse: Text-To-Image Generation via Masked Generative Transformers

Chang, Huiwen; Zhang, Han; Barber, Jarred; Maschinot, AJ; Lezama, Jose; Jiang, Lu; Yang, Ming-Hsuan; Murphy, Kevin; Freeman, William T.; Rubinstein, Michael; Li, Yuanzhen; Krishnan, Dilip

Introduction

Muse learns masked prediction over discrete image tokens with parallel decoding, documenting a text-to-image route distinct from common diffusion sampling and token-by-token autoregression.

ORIGINAL DOCUMENT#210
Figure 1: Muse images with their corresponding prompts.View full image ↗

Figure 1: Muse images with their corresponding prompts. Source: v1; the image version is distinct from the initial submission date.

Chang, Huiwen; Zhang, Han; Barber, Jarred; Maschinot, AJ; Lezama, Jose; Jiang, Lu; Yang, Ming-Hsuan; Murphy, Kevin; Freeman, William T.; Rubinstein, Michael; Li, Yuanzhen; Krishnan, Dilip · original paper / arXiv · Rights in the paper and depicted works remain with their holders. Research quotation does not establish an open licence; republication rights await independent review.

Source · Muse: Text-To-Image Generation via Masked Generative Transformers ↗

RESEARCH ACCOUNT

Muse learns masked prediction over discrete image tokens with parallel decoding, documenting a text-to-image route distinct from common diffusion sampling and token-by-token autoregression.

ANTI-AI ARCHIVE · Revised 2026-10-03

What the paper investigates

A pretrained language model supplies text embeddings while the image model predicts masked tokens. Repeated parallel decoding produces the image. The paper also demonstrates inpainting, outpainting and mask-free editing within this discrete-token framework.

Section sources: Muse: Text-To-Image Generation via Masked Generative Transformers

Its place in generative-art history

Editorial interpretation: recent image-generation history should not treat every system as diffusion. Muse makes discrete representation, parallel generation and editing another line of comparison within the changing conditions of image production.

Section sources: Muse: Text-To-Image Generation via Masked Generative Transformers

Reading limits and versions

This is Google’s 2023 research paper, a different project from the similarly named 2026 Meta product entry.

Section sources: Muse: Text-To-Image Generation via Masked Generative Transformers

Sources for this account

Muse: Text-To-Image Generation via Masked Generative Transformers ↗

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited.

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited. The account and translation are AI-assisted, pending independent human review. Section references identify evidence without claiming independent verification of every historical statement.

Continue with a comparative question

Autoregressive Image Generation without Vector Quantization ↗

Muse: compare methods, control or evaluation conditions with MAR / Diffusion Loss.

Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction ↗

Muse: compare methods, control or evaluation conditions with VAR.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation ↗

Muse: compare methods, control or evaluation conditions with Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis ↗

Muse: compare methods, control or evaluation conditions with PixArt-α.

Dates & version record

Date displayed for this node: 2023-01-02 · Historical date recorded for the source: 2023-01-02

The timeline uses the initial arXiv submission, distinct from conference publication, model release and the archive addition on 3 October 2026. The account and image refer to v1; its date is recorded in the source history.

These dates refer to the historical event or recorded version, not this page’s publication date. The original date precision and unresolved questions are retained.

Original sources & further reading

These links lead to the cited paper, article, institution or conference page. External texts retain their source languages.

01
Muse: Text-To-Image Generation via Masked Generative Transformers ↗

Muse: Text-To-Image Generation via Masked Generative Transformers

https://arxiv.org/abs/2301.00704

Current cited URL

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited.

Provenance, translation & verification

Archive node #210 · Initial research-paper submission; methods, evaluation and production conditions

Research materials & supplement references · 1

Recent research papers: additions and reconciled sources since 2022 · Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited. · 010
Date as recorded: 2023-01-02
Muse: Text-To-Image Generation via Masked Generative Transformers ↗
Muse: Text-To-Image Generation via Masked Generative Transformers

Archive account based on the listed research materials, not a full translation of the linked work. AI-assisted translation; independent human review pending.

Official material was read within the stated scope; see the source note for reading limits and outstanding checks.

Cite this node

ANTI-AI ARCHIVE. Muse: Text-To-Image Generation via Masked Generative Transformers. Art-history node #210. https://salondesrefuses.cn/en/art-history/333

For specific historical claims, also cite the original sources above and include your access date. This account is not a full translation of the linked work.

Adjacent nodes follow chronological order; adjacency does not establish direct influence or causation.