ANTI-AI ARCHIVE
ART-HISTORY NODE / 248 · 2023-10-17

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Ghosh, Dhruba; Hajishirzi, Hanna; Schmidt, Ludwig

Introduction

GenEval decomposes text–image alignment into object co-occurrence, position, count and colour, making “good generation” more specific than a single overall quality score.

ORIGINAL DOCUMENT#248
Figure 1: an evaluation pipeline using object detection, counts and colour checks.View full image ↗

Figure 1: an evaluation pipeline using object detection, counts and colour checks. Source: v1; the image version is distinct from the initial submission date.

Ghosh, Dhruba; Hajishirzi, Hanna; Schmidt, Ludwig · original paper / arXiv · Rights in the paper and depicted works remain with their holders. Research quotation does not establish an open licence; republication rights await independent review.

Source · GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment ↗

RESEARCH ACCOUNT

GenEval decomposes text–image alignment into object co-occurrence, position, count and colour, making “good generation” more specific than a single overall quality score.

ANTI-AI ARCHIVE · Revised 2026-10-03

What the paper investigates

Object detection and other vision models check compositional properties, with comparisons to human judgement. The study identifies difficulties with spatial relations and attribute binding and explains why holistic FID or CLIP scores cannot locate every failure to follow a description.

Section sources: GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Its place in generative-art history

Editorial interpretation: evaluation tools shape which visual capacities count as progress. Measurable object relationships matter to composition, but they cannot replace criticism concerned with ambiguity, symbolism and context.

Section sources: GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reading limits and versions

Detectors have their own errors. The framework measures selected aspects of prompt following, not universal aesthetic or artistic value.

Section sources: GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Sources for this account

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment ↗

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited.

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited. The account and translation are AI-assisted, pending independent human review. Section references identify evidence without claiming independent verification of every historical statement.

Continue with a comparative question

VBench: Comprehensive Benchmark Suite for Video Generative Models ↗

GenEval: compare methods, control or evaluation conditions with VBench.

PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis ↗

GenEval: compare methods, control or evaluation conditions with PixArt-α.

Imagen 3 ↗

GenEval: compare methods, control or evaluation conditions with Imagen 3.

Dates & version record

Date displayed for this node: 2023-10-17 · Historical date recorded for the source: 2023-10-17

The timeline uses the initial arXiv submission, distinct from conference publication, model release and the archive addition on 3 October 2026. The account and image refer to v1; its date is recorded in the source history.

These dates refer to the historical event or recorded version, not this page’s publication date. The original date precision and unresolved questions are retained.

Original sources & further reading

These links lead to the cited paper, article, institution or conference page. External texts retain their source languages.

01
GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment ↗

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

https://arxiv.org/abs/2310.11513

Current cited URL

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited.

Provenance, translation & verification

Archive node #248 · Initial research-paper submission; methods, evaluation and production conditions

Research materials & supplement references · 1

Recent research papers: additions and reconciled sources since 2022 · Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited. · 021
Date as recorded: 2023-10-17
GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment ↗
GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Archive account based on the listed research materials, not a full translation of the linked work. AI-assisted translation; independent human review pending.

Official material was read within the stated scope; see the source note for reading limits and outstanding checks.

Cite this node

ANTI-AI ARCHIVE. GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment. Art-history node #248. https://salondesrefuses.cn/en/art-history/344

For specific historical claims, also cite the original sources above and include your access date. This account is not a full translation of the linked work.

Adjacent nodes follow chronological order; adjacency does not establish direct influence or causation.