ANTI-AI ARCHIVE
ART-HISTORY NODE / 213 · 2023-01-29

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Liu, Haohe; Chen, Zehua; Yuan, Yi; Mei, Xinhao; Liu, Xubo; Mandic, Danilo; Wang, Wenwu; Plumbley, Mark D.

Introduction

AudioLDM applies latent diffusion to general sound synthesis through language–audio representations, extending the research timeline to environmental audio and sound editing.

ORIGINAL DOCUMENT#213
Figure 1: the original AudioLDM diagram of training, text-conditioned sampling and audio editing.View full image ↗

Figure 1: the original AudioLDM diagram of training, text-conditioned sampling and audio editing. Source: v3; the image version is distinct from the initial submission date.

Liu, Haohe; Chen, Zehua; Yuan, Yi; Mei, Xinhao; Liu, Xubo; Mandic, Danilo; Wang, Wenwu; Plumbley, Mark D. · original paper / arXiv · Rights in the paper and depicted works remain with their holders. Research quotation does not establish an open licence; republication rights await independent review.

Source · AudioLDM: Text-to-Audio Generation with Latent Diffusion Models ↗

RESEARCH ACCOUNT

AudioLDM applies latent diffusion to general sound synthesis through language–audio representations, extending the research timeline to environmental audio and sound editing.

ANTI-AI ARCHIVE · Revised 2026-10-03

What the paper investigates

The system uses aligned CLAP representations, training with audio embeddings and conditioning sampling with text embeddings. Its demonstrations also explore audio manipulation, showing that sound generation involves its own representations and transformations rather than merely changing an image model’s output label.

Section sources: AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Its place in generative-art history

Editorial interpretation: text becomes an input for organising sonic environments, bringing timbre, events and imagined spaces into generative production. Author-hosted audio examples remain essential because a static diagram cannot communicate the resulting sound.

Section sources: AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reading limits and versions

General sound synthesis differs from generating musical structure. Objective and listening evaluations do not determine artistic value.

Section sources: AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Sources for this account

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models ↗

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited.

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited. The account and translation are AI-assisted, pending independent human review. Section references identify evidence without claiming independent verification of every historical statement.

Continue with a comparative question

Simple and Controllable Music Generation ↗

AudioLDM: compare methods, control or evaluation conditions with MusicGen.

Jukebox: A Generative Model for Music ↗

AudioLDM: compare methods, control or evaluation conditions with Jukebox.

IMPF and IMPEL publish principles for fair generative-AI licensing ↗

AudioLDM: compare methods, control or evaluation conditions with Fair music-AI licensing.

Dates & version record

Date displayed for this node: 2023-01-29 · Historical date recorded for the source: 2023-01-29

The timeline uses the initial arXiv submission, distinct from conference publication, model release and the archive addition on 3 October 2026. The account and image refer to v3; its date is recorded in the source history.

These dates refer to the historical event or recorded version, not this page’s publication date. The original date precision and unresolved questions are retained.

Original sources & further reading

These links lead to the cited paper, article, institution or conference page. External texts retain their source languages.

01
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models ↗

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

https://arxiv.org/abs/2301.12503

Current cited URL

Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited.

Provenance, translation & verification

Archive node #213 · Initial research-paper submission; methods, evaluation and production conditions

Research materials & supplement references · 1

Recent research papers: additions and reconciled sources since 2022 · Read the arXiv abstract, authors and version history, plus the selected figure/page on PDF page 2; proofs and full experiments were not independently audited. · 012
Date as recorded: 2023-01-29
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models ↗
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Archive account based on the listed research materials, not a full translation of the linked work. AI-assisted translation; independent human review pending.

Official material was read within the stated scope; see the source note for reading limits and outstanding checks.

Cite this node

ANTI-AI ARCHIVE. AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. Art-history node #213. https://salondesrefuses.cn/en/art-history/335

For specific historical claims, also cite the original sources above and include your access date. This account is not a full translation of the linked work.

Adjacent nodes follow chronological order; adjacency does not establish direct influence or causation.