This guide connects 52 research-related entries: 33 papers from 2022–2025 added in this pass, five existing achievements enriched with paper and version sources, and previously catalogued research. Seven paths cover architectures and sampling, editing, 3D, video, sound, evaluation and cultural inquiry. Selection prioritises distinct methodological contributions, production and viewing conditions, and explanatory value for evaluation, data and artists’ rights, rather than companies or product popularity. Initial preprint submission supplies the new timeline date; conference versions, revisions and releases remain separate. Figures retain source, credit and version context, with demonstrations linked to author projects. This pass reads abstracts, submission records and specified figure pages without reproducing experiments. It does not exhaust 2022–2026 scholarship; the longer-term significance of recent 2026 candidates remains under review.
Follow this reading path
Architectures, sampling and training
2022 · Hierarchical Text-Conditional Image Generation with CLIP Latents / DALL·E 2 product ↗DALL·E 2 combines CLIP representations with a diffusion-based image prior and decoder, improving realism and correspondence with text. Its research preview, beta and later access without a waitlist expand during 2022. The master emphasises this product transition as a step in the international popularization of text-to-image generation.
2022–2023 · LoRA: Low-Rank Adaptation of Large Language Models + diffusion community adoption ↗The original LoRA paper appears in 2021, reducing fine-tuning costs by freezing a main model and training low-rank update matrices. During 2022–2023, the Stable Diffusion community adapts this into small exchangeable modules for styles, characters, clothing, poses and likenesses. Style and identity can circulate as lightweight model files as well as names in prompts.
2022-05 · Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding ↗Imagen uses a large text encoder and diffusion decoders. The research highlights gains in image–text alignment from scaling the language model. It develops a direction that later multimodal systems also pursue: improvements in understanding language can change the visual capabilities and controllability of image generation.
2022-06 · Scaling Autoregressive Models for Content-Rich Text-to-Image Generation ↗Parti explores an autoregressive approach alongside diffusion, generating image tokens as another kind of language. It shows that the expansion of text-to-image generation in 2022 did not follow a single technical route. Token-based modelling remains a parallel way to organize the relation between descriptions and visual outputs.
2022-06-01 · Elucidating the Design Space of Diffusion-Based Generative Models ↗EDM separates sampling, training and network preconditioning into comparable design choices, making speed and image quality questions of method rather than model branding.
2022-06-02 · DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps ↗DPM-Solver designs a dedicated solver for diffusion ordinary differential equations, reducing network evaluations and adding a history of sampling methods behind model releases.
2022-09-07 · Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow ↗Rectified Flow formulates generation and distribution transfer through increasingly straight transport paths, supplying a theoretical source for later flow-based image models independently of product announcements.
2022-10-06 · Flow Matching for Generative Modeling ↗Flow Matching trains continuous flow models by regressing vector fields along conditional probability paths, bringing diffusion and alternative transport paths into a shared methodological framework.
2022-12-19 · Scalable Diffusion Models with Transformers ↗DiT replaces the usual U-Net backbone with a transformer operating on latent patches and studies compute scaling, providing a key architectural comparison for later generative models.
2023-01-02 · Muse: Text-To-Image Generation via Masked Generative Transformers ↗Muse learns masked prediction over discrete image tokens with parallel decoding, documenting a text-to-image route distinct from common diffusion sampling and token-by-token autoregression.
2023-03-02 · Consistency Models ↗Consistency Models investigates one- and few-step generation through either distillation or standalone training, supplying a methodological background for later interactive image-generation workflows.
2023-07-26 · SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis ↗SDXL develops composition, detail, style and resolution while retaining a model ecosystem that can be fine-tuned and combined with other tools. The master positions it within the Stable Diffusion community's movement from rapid experimentation toward more established production practices.
2023-09 · PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis ↗PixArt-α focuses on diffusion Transformers, pretrained text encoders and efficient training. It is an open research node leading into the wider development of Transformer- and flow-based image systems in 2024. Its significance concerns both architectural choices and the resources needed to train visual models.
2023-10 · Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference ↗Latent Consistency Models use consistency distillation in latent space to generate high-resolution images in a small number of steps, including two to four in the reported approach. Speed changes the possible interface: generation closer to real time can participate more readily in drawing, design, live presentation and interactive software.
2024-02 · Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis ↗Stable Diffusion 3 combines a Transformer-based architecture with flow matching and emphasises multi-subject prompts, text and image quality. It represents the further incorporation of Transformer approaches into mainstream image generation. The associated research paper provides a technical account distinct from the product announcement.
2024-04-03 · Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction ↗VAR reformulates visual autoregression as coarse-to-fine next-scale prediction, providing a distinct image-generation route to compare with diffusion transformers.
2024-06-17 · Autoregressive Image Generation without Vector Quantization ↗The paper models continuous-token distributions with diffusion, connecting autoregression to continuous latents and showing that autoregressive image generation need not require discrete quantisation.
2024-10-09 · Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think ↗REPA aligns intermediate denoising representations with pretrained visual encoders, investigating how visual representations can support diffusion and flow-transformer training.
2024-10-14 · SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers ↗SANA combines deep latent compression, linear attention and efficient sampling for high-resolution synthesis, connecting model design with computation and local production conditions.
2025-01-29 · Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling ↗Janus-Pro studies training, data and model scaling in a unified multimodal system, providing follow-up evidence for the Janus approach to understanding and image generation.
2025-05-19 · Mean Flows for One-step Generative Modeling ↗MeanFlow characterises generative flows through interval-average rather than instantaneous velocity and trains one-step models from scratch, extending the history of few-step generation.
Image editing and control
2022-08-02 · An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion ↗Textual Inversion learns a new text embedding without retraining an entire model. Visual concepts, objects and styles can be represented through a new composable word within the system. This is an effective method of personalisation, not a claim that the nature of artistic style can be reduced to a token.
2022-08-02 · Prompt-to-Prompt Image Editing with Cross Attention Control ↗Prompt-to-Prompt studies cross-attention between words and image layout, using prompt changes for local replacement and broader transformation. It connects image synthesis with the problem of editable continuity.
2022-08-25 · DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation ↗DreamBooth fine-tunes a pretrained text-to-image model using a small set of subject images, associating a unique identifier with that subject. It brings preservation of identity into the generative process and supplies technical context for later customization of people, characters and products, as well as disputes concerning likeness.
2022-11 / 2023 · InstructPix2Pix: Learning to Follow Image Editing Instructions ↗InstructPix2Pix uses GPT-3 and Stable Diffusion to synthesize training examples, then trains a diffusion model conditioned on editing instructions. It shifts the interaction from making a picture in a single generation toward changing an existing image through language, connecting generative modelling with iterative visual editing.
2022-11-17 · Null-text Inversion for Editing Real Images using Guided Diffusion Models ↗Null-text Inversion optimises unconditional text embeddings to invert real images, connecting prompt-based editing to existing photographs without retraining the model weights.
2023-01-17 · GLIGEN: Open-Set Grounded Text-to-Image Generation ↗GLIGEN adds gated layers to inject grounding conditions into a pretrained diffusion model, using text and bounding boxes to control objects and layout beyond a prompt alone.
2023-02 · Adding Conditional Control to Text-to-Image Diffusion Models ↗ControlNet freezes a pretrained diffusion backbone and learns spatial conditions through an additional network. It extends control from textual semantics toward geometry and composition. This is one technical condition for more specialized production workflows, in which artists or designers need to specify structure as well as subject matter.
2023-02 · T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models ↗Alongside ControlNet, T2I-Adapter exemplifies a move toward fine-grained external control through modular attachments without fully retraining a foundation model. Image production becomes a combination of a base model and adapters. The node situates controllability in a growing ecosystem of interoperable components.
2023-05 · Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold ↗DragGAN uses a GAN's feature space for point-based manipulation, demonstrating direct interaction with a generative image manifold. Although diffusion becomes increasingly prominent, this approach remains relevant to the development of editing interfaces in which users move elements visually instead of specifying every change through text.
2023-08 · IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models ↗IP-Adapter uses decoupled cross-attention to add lightweight image-prompt capabilities to pretrained diffusion models. Reference-image control can coexist with text prompts and ControlNet, supporting workflows for portraits, fashion and identity that combine several types of guidance rather than relying on a description alone.
2024-01 · InstantID: Zero-shot Identity-Preserving Generation in Seconds ↗InstantID combines facial embeddings with structural conditions to support efficient identity preservation in diffusion models. It marks a move from training specifically on a person, as in DreamBooth-style workflows, toward generation guided by references. The technique is directly relevant to later questions of likeness and consent.
3D generation and spatial representations
2022-09-29 · DreamFusion: Text-to-3D using 2D Diffusion ↗DreamFusion uses a two-dimensional diffusion prior to optimise a three-dimensional representation, extending text-driven synthesis toward objects that can be viewed and lit from different directions.
2023-03-20 · Zero-1-to-3: Zero-shot One Image to 3D Object ↗Zero-1-to-3 conditions novel views on a single image and a relative camera change, connecting image priors with 3D reconstruction while exposing the problem of inferred unseen surfaces.
2023-05-25 · ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation ↗ProlificDreamer introduces variational score distillation to address oversaturation, oversmoothing and limited diversity, documenting a methodological change in text-to-3D generation.
2023-08-08 · 3D Gaussian Splatting for Real-Time Radiance Field Rendering ↗3D Gaussian Splatting represents scenes with optimisable 3D Gaussians and renders new views efficiently. It is spatial-image infrastructure, not itself a text-to-3D generator.
2024-12-02 · Structured 3D Latents for Scalable and Versatile 3D Generation ↗TRELLIS uses structured latents to connect text or image conditions with meshes, radiance fields and Gaussians, examining representation and editability in 3D asset generation.
Video and temporal continuity
2022-04-07 · Video Diffusion Models ↗Video Diffusion Models extends image denoising into time, combining image and video training with clip extension. It supplies a research account of temporal continuity behind later video-generation services.
2023 · AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning ↗AnimateDiff introduces motion modelling as a pluggable component within personalized text-to-image diffusion systems. Existing character models, styles and community assets can participate in animated outputs. This extends generation from still images toward video while retaining parts of the earlier image-model ecosystem.
2023-11 · Stable Video Diffusion ↗Stable Video Diffusion makes latent video diffusion available to researchers and developers, extending an open generative ecosystem from still images toward video. It also points toward increasingly shared infrastructure between image and video models, including software for arranging and adapting generation processes.
2024-01-23 · Lumiere: A Space-Time Diffusion Model for Video Generation ↗Lumiere uses a space–time U-Net that processes the full video duration in a model pass, offering an alternative to sparse keyframes followed by temporal interpolation.
2025-04-17 · Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models ↗FramePack packs past-frame context by importance and studies generation drift, making memory and computation in longer video generation an independent research entry.
2025-06-09 · Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion ↗Self Forcing trains video models on their own generated context, addressing the mismatch between ground-truth training frames and self-generated inference history in streaming generation.
Sound and music generation
2023-01-29 · AudioLDM: Text-to-Audio Generation with Latent Diffusion Models ↗AudioLDM applies latent diffusion to general sound synthesis through language–audio representations, extending the research timeline to environmental audio and sound editing.
2023-06-08 · Simple and Controllable Music Generation ↗MusicGen organises multiple streams of compressed music tokens through a single language model, using text or melody conditions and adding sonic representation and control to the timeline.
Evaluation and benchmarks
2023-10-17 · GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment ↗GenEval decomposes text–image alignment into object co-occurrence, position, count and colour, making “good generation” more specific than a single overall quality score.
2023-11-29 · VBench: Comprehensive Benchmark Suite for Video Generative Models ↗VBench separates video quality into 16 dimensions, including subject consistency, motion, flicker and spatial relations, and checks metrics against human annotations.
Data, artists’ rights and cultural inquiry
2022 · LAION-5B: An open large-scale dataset for training next generation image-text models ↗LAION-5B makes a large image–text pair index and filtering methods available, providing data infrastructure relevant to Stable Diffusion and related ecosystems. It also becomes central to disputes about training sources, artists' names and copyright. The master reads this as a shift in which the dataset itself becomes a political object; an index does not confer ownership of linked images.
2023-01-30 · Extracting Training Data from Diffusion Models ↗This paper tests memorisation through a generate-and-filter procedure, making training-image reproduction an empirical research question rather than an inference from visual resemblance alone.
2023-02-08 · Glaze: Protecting Artists from Style Mimicry by Text-to-Image Models ↗Glaze studies small image perturbations intended to disrupt style-mimicking fine-tuning and includes artist studies, making refusal of model learning part of the technical history of generative art.
2023-06-07 · Art and the science of generative AI: A deeper dive ↗Epstein and colleagues outline a research agenda spanning aesthetics and culture, ownership and credit, creative labour and media ecosystems, extending discussion beyond model performance.
2023-10-20 · Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models ↗Nightshade studies prompt-specific training-data poisoning and discusses creator resistance to scraping that ignores opt-out requests, connecting technical research with disputes over data control.
Reference archives and reading scope
38 individual paper records checked in this pass; each entry specifies its reading scope. This is not a claim of a complete arXiv search.
An archive research guide. Chinese and English text may include AI-assisted translation and awaits independent human review. Check historical claims against each node’s original sources.
Continue reading
Generative Art History: historical threads, 1943–2026 ↗Why begin in 1943? ↗How can you use these 356 nodes? ↗Generative-art timeline: early computing, software art and crypto ecosystems ↗Return to the art-history catalogue ↗Cite this page
ANTI-AI ARCHIVE. Important papers since 2022: methods and cultural questions. https://salondesrefuses.cn/en/art-history/guides/recent-papers
Include your access date. For historical claims, also cite the relevant node and original source.

