ANTI-AI ARCHIVE
ART-HISTORY NODE / 311 · 2025-01-29

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Chen, Xiaokang; Wu, Zhiyu; Liu, Xingchao; Pan, Zizheng; Liu, Wen; Xie, Zhenda; Yu, Xingkai; Ruan, Chong

Introduction

Janus-Pro studies training, data and model scaling in a unified multimodal system, providing follow-up evidence for the Janus approach to understanding and image generation.

ORIGINAL DOCUMENT#311
Figure 1: the authors’ reported understanding and generation benchmark comparisons for Janus-Pro.View full image ↗

Figure 1: the authors’ reported understanding and generation benchmark comparisons for Janus-Pro. Source: v1; the image version is distinct from the initial submission date.

Chen, Xiaokang; Wu, Zhiyu; Liu, Xingchao; Pan, Zizheng; Liu, Wen; Xie, Zhenda; Yu, Xingkai; Ruan, Chong · original paper / arXiv · Rights in the paper and depicted works remain with their holders. Research quotation does not establish an open licence; republication rights await independent review.

Source · Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling ↗

RESEARCH ACCOUNT

Janus-Pro studies training, data and model scaling in a unified multimodal system, providing follow-up evidence for the Janus approach to understanding and image generation.

ANTI-AI ARCHIVE · Revised 2026-10-03

What the paper investigates

The report separates changes in training strategy, expanded data and model scale, comparing multimodal understanding with text-to-image instruction following. It explicitly develops the prior Janus approach rather than originating every idea of multimodal unification.

Section sources: Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Its place in generative-art history

Editorial interpretation: bringing image understanding and production into one system can support interaction around an image rather than a single generation request. This is a background for interface history; evidence of artists’ use remains a separate research task.

Section sources: Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reading limits and versions

The entry uses the 29 January 2025 preprint date, distinct from code or model releases; benchmark gains do not establish general visual reasoning.

Section sources: Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Sources for this account

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling ↗

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. The account and translation are AI-assisted, pending independent human review. Section references identify evidence without claiming independent verification of every historical statement.

Continue with a comparative question

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment ↗

Janus-Pro: compare methods, control or evaluation conditions with GenEval.

Gemini native image generation preview ↗

Janus-Pro: compare methods, control or evaluation conditions with Gemini native image generation preview.

GPT-4o native image generation ↗

Janus-Pro: compare methods, control or evaluation conditions with Images in conversation.

Dates & version record

Date displayed for this node: 2025-01-29 · Historical date recorded for the source: 2025-01-29

The timeline uses the initial arXiv submission, distinct from conference publication, model release and the archive addition on 3 October 2026. The account and image refer to v1; its date is recorded in the source history.

These dates refer to the historical event or recorded version, not this page’s publication date. The original date precision and unresolved questions are retained.

Original sources & further reading

These links lead to the cited paper, article, institution or conference page. External texts retain their source languages.

01
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling ↗

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

https://arxiv.org/abs/2501.17811

Current cited URL

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Provenance, translation & verification

Archive node #311 · Initial research-paper submission; methods, evaluation and production conditions

Research materials & supplement references · 1

Recent research papers: additions and reconciled sources since 2022 · Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. · 030
Date as recorded: 2025-01-29
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling ↗
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Archive account based on the listed research materials, not a full translation of the linked work. AI-assisted translation; independent human review pending.

Official material was read within the stated scope; see the source note for reading limits and outstanding checks.

Cite this node

ANTI-AI ARCHIVE. Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling. Art-history node #311. https://salondesrefuses.cn/en/art-history/353

For specific historical claims, also cite the original sources above and include your access date. This account is not a full translation of the linked work.

Adjacent nodes follow chronological order; adjacency does not establish direct influence or causation.