ANTI-AI ARCHIVE
ART-HISTORY NODE / 211 · 2023-01-17

GLIGEN: Open-Set Grounded Text-to-Image Generation

GLIGEN: Open-Set Grounded Text-to-Image Generation

Li, Yuheng; Liu, Haotian; Wu, Qingyang; Mu, Fangzhou; Yang, Jianwei; Gao, Jianfeng; Li, Chunyuan; Lee, Yong Jae

Introduction

GLIGEN adds gated layers to inject grounding conditions into a pretrained diffusion model, using text and bounding boxes to control objects and layout beyond a prompt alone.

ORIGINAL DOCUMENT#211
Figure 1: generation with boxes, keypoints and other spatial conditions.View full image ↗

Figure 1: generation with boxes, keypoints and other spatial conditions. Source: v2; the image version is distinct from the initial submission date.

Li, Yuheng; Liu, Haotian; Wu, Qingyang; Mu, Fangzhou; Yang, Jianwei; Gao, Jianfeng; Li, Chunyuan; Lee, Yong Jae · original paper / arXiv · Rights in the paper and depicted works remain with their holders. Research quotation does not establish an open licence; republication rights await independent review.

Source · GLIGEN: Open-Set Grounded Text-to-Image Generation ↗

RESEARCH ACCOUNT

GLIGEN adds gated layers to inject grounding conditions into a pretrained diffusion model, using text and bounding boxes to control objects and layout beyond a prompt alone.

ANTI-AI ARCHIVE · Revised 2026-10-03

What the paper investigates

The original model weights stay frozen while new trainable layers receive grounding information through gates. Captions and bounding boxes provide conditions. Experiments test how this control generalises to unfamiliar concepts and spatial arrangements.

Section sources: GLIGEN: Open-Set Grounded Text-to-Image Generation

Its place in generative-art history

Editorial interpretation: bounding boxes make “where” an explicit compositional input. Comparison with ControlNet should distinguish conditioning types and control mechanisms, rather than presenting all controllable generation as a single interchangeable tool.

Section sources: GLIGEN: Open-Set Grounded Text-to-Image Generation

Reading limits and versions

Successful grounding experiments do not guarantee exact execution of every object and relationship in a complex scene.

Section sources: GLIGEN: Open-Set Grounded Text-to-Image Generation

Sources for this account

GLIGEN: Open-Set Grounded Text-to-Image Generation ↗

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. The account and translation are AI-assisted, pending independent human review. Section references identify evidence without claiming independent verification of every historical statement.

Continue with a comparative question

Prompt-to-Prompt Image Editing with Cross Attention Control ↗

GLIGEN: compare methods, control or evaluation conditions with Prompt-to-Prompt.

Adding Conditional Control to Text-to-Image Diffusion Models ↗

GLIGEN: compare methods, control or evaluation conditions with Spatial control of generation.

T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models ↗

GLIGEN: compare methods, control or evaluation conditions with T2I-Adapter.

Dates & version record

Date displayed for this node: 2023-01-17 · Historical date recorded for the source: 2023-01-17

The timeline uses the initial arXiv submission, distinct from conference publication, model release and the archive addition on 3 October 2026. The account and image refer to v2; its date is recorded in the source history.

These dates refer to the historical event or recorded version, not this page’s publication date. The original date precision and unresolved questions are retained.

Original sources & further reading

These links lead to the cited paper, article, institution or conference page. External texts retain their source languages.

01
GLIGEN: Open-Set Grounded Text-to-Image Generation ↗

GLIGEN: Open-Set Grounded Text-to-Image Generation

https://arxiv.org/abs/2301.07093

Current cited URL

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited.

Provenance, translation & verification

Archive node #211 · Initial research-paper submission; methods, evaluation and production conditions

Research materials & supplement references · 1

Recent research papers: additions and reconciled sources since 2022 · Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 1; proofs and full experiments were not independently audited. · 011
Date as recorded: 2023-01-17
GLIGEN: Open-Set Grounded Text-to-Image Generation ↗
GLIGEN: Open-Set Grounded Text-to-Image Generation

Archive account based on the listed research materials, not a full translation of the linked work. AI-assisted translation; independent human review pending.

Official material was read within the stated scope; see the source note for reading limits and outstanding checks.

Cite this node

ANTI-AI ARCHIVE. GLIGEN: Open-Set Grounded Text-to-Image Generation. Art-history node #211. https://salondesrefuses.cn/en/art-history/334

For specific historical claims, also cite the original sources above and include your access date. This account is not a full translation of the linked work.

Adjacent nodes follow chronological order; adjacency does not establish direct influence or causation.