AI art history is more than a sequence of image-model releases. This reading path brings theory, artists’ practices, exhibition institutions and technical tools together to ask who sets the rules, who participates in making a work, and how that work is exhibited and interpreted.
Why begin in 1943?
1943 is an intellectual-context marker selected by the research master, not a claim that AI art began in that year. Early neural computation, cybernetics, computer art and later machine learning are not interchangeable. Their differences matter; this genealogy does not treat them as a single, inevitable progression toward present-day generative AI.
How can you use these 140 nodes?
Start with a period, then compare original titles, people and institutions, research accounts and source links. A theoretical paper can document a proposed method; an exhibition record can document display and circulation. Neither alone establishes a technology’s influence on all artistic practice. Each entry retains date precision, source limitations and verification status. The sidebar offers interactive browsing; the text below supports continuous reading and citation.
This archive-written research guide includes AI-assisted Chinese and English texts awaiting independent human review. Check specific historical claims and rights information against each record’s original-language sources.
140 historical nodes
Expand an entry to read its research account and source. Identifiers follow the research master; entries are ordered chronologically. Historical claims await independent verification.
1943A Logical Calculus of the Ideas Immanent in Nervous Activity
+
A Logical Calculus of the Ideas Immanent in Nervous Activity
McCulloch & Pitts
McCulloch and Pitts use simplified neurons and logical calculus to examine how neural activity could perform computation. This is not a paper on image generation. Its importance to this genealogy is the proposition that perception and cognition can be represented as computable networks: a formal prehistory that should remain visible when discussing neural image generation.
Read original source · A Logical Calculus of the Ideas Immanent in Nervous Activity ↗
Research master · p. 4 · #001 · Research claims await independent verification
Open in the interactive timeline ↗1948A Mathematical Theory of Communication
+
A Mathematical Theory of Communication
Claude Shannon
Shannon establishes a mathematical theory of communication. The genealogy does not position him as the direct origin of AI art; it identifies a broader condition in which images, writing and sound can be treated as encodable information. Datasets, compressed representations, latent spaces and generative models develop within this wider informational framework.
Read original source · A Mathematical Theory of Communication ↗
Research master · p. 4 · #002 · Research claims await independent verification
Open in the interactive timeline ↗1948Cybernetics
+
Cybernetics
Norbert Wiener
Wiener's Cybernetics brings communication, feedback and control into one theoretical framework. For the archive's later questions of control and sovereignty, cultural information is not merely stored or transmitted: it can enter feedback loops and become a resource for regulating machine behavior and organizing power. This connection is the master's historical interpretation.
Read original source · Cybernetics ↗
Research master · p. 4 · #003 · Research claims await independent verification
Open in the interactive timeline ↗1949The Organization of Behavior
+
The Organization of Behavior
Donald Hebb
Hebb's account of learning emphasizes that connections between neurons can change through joint activity. Although far removed from image generation, it helps establish a consequential hypothesis: capabilities need not be written into a system entirely as explicit rules, but may be acquired through experience and changes in connection strengths.
Read original source · The Organization of Behavior ↗
Research master · p. 5 · #004 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗1950Computing Machinery and Intelligence
+
Computing Machinery and Intelligence
Alan Turing
Rather than first defining the essence of thought, Turing's imitation game reframes machine intelligence as a question of behavior and interaction. The master connects this operational shift with later assessments of visual generation through what systems can produce, modify and organize, rather than through a prior resolution of whether machines possess creativity.
Read original source · Computing Machinery and Intelligence ↗
Research master · p. 5 · #005 · Research claims await independent verification
Open in the interactive timeline ↗1952Oscillons / Electronic Abstractions
+
Oscillons / Electronic Abstractions
Ben F. Laposky
Laposky generates electronic waveforms with an oscilloscope and records them photographically as abstract images. Before digital computer art becomes established, this practice connects electrical signals, mathematical curves and visual composition. It forms a prehistory of electronic generative aesthetics rather than an instance of contemporary machine learning.
Read original source · Oscillons / Electronic Abstractions ↗
Research master · p. 7 · #012 · Research claims await independent verification
Open in the interactive timeline ↗1954–1969Aesthetica / Introduction to Information-Theoretic Aesthetics
+
Aesthetica / Einführung in die informationstheoretische Ästhetik
Max Bense
Bense publishes four volumes of Aesthetica in 1954–1960, followed by an expanded collected edition in 1965 and an introduction to information-theoretic aesthetics in 1969. These works connect aesthetics, information theory, statistical structure and programmed production. The genealogy marks a turn toward computational descriptions of form, order, complexity and aesthetic difference; it does not assert that art is simply information.
Read original source · Aesthetica / Einführung in die informationstheoretische Ästhetik ↗
Research master · p. 5 · #006 · Research claims await independent verification
Open in the interactive timeline ↗1956Dartmouth Summer Research Project on Artificial Intelligence proposal/workshop
+
Dartmouth Summer Research Project on Artificial Intelligence proposal/workshop
John McCarthy; Marvin Minsky; Nathaniel Rochester; Claude Shannon and participants
The Dartmouth workshop does not directly concern art. It establishes an academic community around the proposition that intelligence can be simulated and studied by machines. Later computer vision, machine learning and generative models develop within this institutional lineage. The node therefore provides disciplinary context for art history rather than documenting an artistic event.
Research master · p. 5 · #007 · Research claims await independent verification
Open in the interactive timeline ↗1956–1959Oscillograms; early computer / electronic art writing
+
Oscillograms; early computer / electronic art writing
Herbert W. Franke
Franke makes electronic oscillographic images and writes extensively about the roles of mathematics, information and technology in art. Alongside the information-aesthetic milieu associated with Bense and Nake, his work helps establish networks of theory and practice in early computer art in Germany and Austria.
Read original source · Oscillograms; early computer / electronic art writing ↗
Research master · p. 7 · #013 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗1958The Perceptron
+
The Perceptron
Frank Rosenblatt
Rosenblatt's perceptron links pattern recognition to neural-network learning. In the history of visual AI, it marks an early step from systems whose rules are entirely specified by programmers toward systems that adjust their parameters using examples. Its relevance here concerns this change in how visual capabilities are constructed.
Read original source · The Perceptron ↗
Research master · p. 6 · #008 · Research claims await independent verification
Open in the interactive timeline ↗1958 / 1966Information Theory and Esthetic Perception
+
Théorie de l'information et perception esthétique
Abraham Moles
Moles's French study later appears in English as Information Theory and Esthetic Perception. It examines how art objects are perceived as informational messages and how redundancy and complexity contribute to aesthetic experience. Alongside Bense's work, it supplies an important theoretical background for information aesthetics and subsequent debates about computational form.
Read original source · Théorie de l'information et perception esthétique ↗
Research master · p. 6 · #009 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗1959Some Studies in Machine Learning Using the Game of Checkers
+
Some Studies in Machine Learning Using the Game of Checkers
Arthur Samuel
Samuel's checkers program demonstrates how a machine can improve its performance through learning without a programmer specifying every strategy step by step. The master situates this shift from explicit rules toward learning from experience within the longer technical history of data-driven visual systems and contemporary generative models.
Read original source · Some Studies in Machine Learning Using the Game of Checkers ↗
Research master · p. 6 · #010 · Research claims await independent verification
Open in the interactive timeline ↗1960Man-Computer Symbiosis
+
Man-Computer Symbiosis
J. C. R. Licklider
Licklider proposes human–computer symbiosis: people and computers undertake different tasks and form a new cognitive arrangement through real-time interaction. The master relates this history of interaction to contemporary conversational image generation and visual agents. The connection concerns interfaces for collaborative work, rather than an identity between the earlier proposal and present models.
Read original source · Man-Computer Symbiosis ↗
Research master · p. 6 · #011 · Research claims await independent verification
Open in the interactive timeline ↗1961–1973New Tendencies exhibitions and Computers and Visual Research
+
New Tendencies exhibitions and Computers and Visual Research
Almir Mavignier · Matko Meštrović · international participants
New Tendencies in Zagreb moves from concrete and optical art toward Computers and Visual Research, placing programs, information, scientific methods and viewers' perception at the center of artistic discussion. The series provides an early international institutional platform for computer art, extending the genealogy beyond individual inventions or isolated laboratories.
Read original source · New Tendencies exhibitions and Computers and Visual Research ↗
Research master · p. 7 · #015 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗1963Sketchpad
+
Sketchpad
Ivan Sutherland
Sketchpad brings graphical objects, constraints and human–computer interaction into one system, helping establish interactive computer graphics. It reminds us that a history of AI art concerns not only how models generate images, but also how interfaces turn computational capacities into usable visual practices.
Read original source · Sketchpad ↗
Research master · p. 8 · #016 · Research claims await independent verification
Open in the interactive timeline ↗1964–1966Human or Machine: A Subjective Comparison of Piet Mondrian’s Composition with Lines and a Computer-Generated Picture
+
Human or Machine: A Subjective Comparison of Piet Mondrian’s Composition with Lines and a Computer-Generated Picture
A. Michael Noll
Noll generates computer images resembling a Mondrian composition and uses viewer experiments to compare aesthetic preferences and judgments of authorship. This anticipates later debates about machine imitation of style and the recognition of AI images. The technical mechanism, however, remains programmed generation rather than learning from training images.
Research master · p. 8 · #017 · Research claims await independent verification
Open in the interactive timeline ↗1965Early computer-generated art exhibitions and Noll experiments
+
Early computer-generated art exhibitions and Noll experiments
Georg Nees; Frieder Nake; A. Michael Noll
In 1965, Nees exhibits computer graphics in Stuttgart; Noll and Béla Julesz show computer-generated pictures at New York's Howard Wise Gallery; Nake is also among the earliest public exhibitors of algorithmic art. The computer enters art institutions as a means of producing form, extending its role beyond engineering.
Read original source · Early computer-generated art exhibitions and Noll experiments ↗
Research master · p. 8 · #018 · Research claims await independent verification
Open in the interactive timeline ↗19669 Evenings: Theatre & Engineering
+
9 Evenings: Theatre & Engineering
Billy Klüver; Robert Rauschenberg; John Cage; Lucinda Childs; engineers and artists
9 Evenings organizes artists and Bell Labs engineers into large-scale live experiments and contributes directly to the formation of E.A.T. The event expands the roles of computation and control systems in art, providing context for later interactive art, machine behavior and real-time generation.
Read original source · 9 Evenings: Theatre & Engineering ↗
Research master · p. 8 · #019 · Research claims await independent verification
Open in the interactive timeline ↗1966Experiments in Art and Technology (organization)
+
Experiments in Art and Technology (organization)
Billy Klüver; Fred Waldhauer; Robert Rauschenberg; Robert Whitman
Billy Klüver, Fred Waldhauer, Robert Rauschenberg and Robert Whitman help establish Experiments in Art and Technology, institutionalizing collaboration between artists and engineers. E.A.T. is not an AI art organization, but its collaborative production model becomes important to subsequent digital, media and experimental technological art.
Read original source · Experiments in Art and Technology (organization) ↗
Research master · p. 9 · #020 · Research claims await independent verification
Open in the interactive timeline ↗1968Computer Arts Society
+
Computer Arts Society
Alan Sutcliffe; George Mallen; John Lansdown
Founded in Britain, the Computer Arts Society connects artists, designers, engineers and researchers. Its significance is institutional: computer art begins to acquire sustained structures for exhibition, discussion, publication and professional exchange, moving beyond scattered experiments into a continuing field of practice.
Read original source · Computer Arts Society ↗
Research master · p. 9 · #021 · Research claims await independent verification
Open in the interactive timeline ↗1968Cybernetic Serendipity exhibition
+
Cybernetic Serendipity exhibition
Jasia Reichardt (curator); participating artists/researchers
Curated by Jasia Reichardt, Cybernetic Serendipity brings together computer music, computer graphics, machines and environments, and computer poetry. It moves the question of computers and art from laboratories into a broader public cultural setting, while placing cybernetics and artistic practice alongside one another.
Read original source · Cybernetic Serendipity exhibition ↗
Research master · p. 9 · #022 · Research claims await independent verification
Open in the interactive timeline ↗1968–1980sAlgorithmic drawings / “machine imaginaire” practice
+
Algorithmic drawings / “machine imaginaire” practice
Vera Molnár
From a manually conceived machine imaginaire to actual computer programs, Molnár uses rules, permutations and small random deviations in painting and drawing. Her practice demonstrates that generative art was not suddenly invented by deep learning: artists had long treated procedures, constraints and variation as mechanisms through which works take form.
Read original source · Algorithmic drawings / “machine imaginaire” practice ↗
Research master · p. 9 · #023 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗late 1960s–1970sAARON
+
AARON
Harold Cohen
Cohen begins conceiving AARON in the late 1960s and names and develops it during the 1970s. Instead of learning from a massive image dataset, AARON encodes his knowledge of drawing, representation and composition as rules. It offers an important comparison: machine art need not depend on scraping or style appropriation, so contemporary anti-AI politics cannot simply be equated with opposition to all machine-made art.
Read original source · AARON ↗
Research master · p. 7 · #014 · Research claims await independent verification
Open in the interactive timeline ↗1969–1971Computer-generated Drawings / A Programmed Aesthetics
+
Computer-generated drawings; Une esthétique programmée
Manfred Mohr
Mohr systematically uses computers and plotters after 1969 and holds a significant solo exhibition at the Musée d'Art Moderne de la Ville de Paris in 1971. His work demonstrates how a program can become an enduring system of authorship, comparable in its continuity to an artist's sustained method of painting.
Read original source · Computer-generated drawings; Une esthétique programmée ↗
Research master · p. 10 · #024 · Research claims await independent verification
Open in the interactive timeline ↗1970Software exhibition, Jewish Museum, New York
+
Software exhibition, Jewish Museum, New York
Jack Burnham (curator); artists and technologists
Jack Burnham's Software exhibition brings information processing, conceptual systems and interactive processes into an art institution. It connects systems aesthetics, conceptual art and computational culture, providing an important context for understanding artworks through rules, protocols, programs and informational processes rather than exclusively as discrete objects.
Read original source · Software exhibition, Jewish Museum, New York ↗
Research master · p. 10 · #025 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗1971Art and Technology program/exhibition
+
Art and Technology program/exhibition
Maurice Tuchman (curator); participating artists and corporations
LACMA's Art and Technology program gives artists access to corporate laboratories and industrial resources. It is not an AI art project. Its place in this genealogy is institutional: it offers an early model for later production relationships among artists, large technology companies and research organizations.
Read original source · Art and Technology program/exhibition ↗
Research master · p. 10 · #026 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗1975 / 1982Fractal Objects (1975) / The Fractal Geometry of Nature (1982)
+
Les objets fractals (1975); The Fractal Geometry of Nature (1982)
Benoît Mandelbrot
Mandelbrot's fractal research is not AI, but has substantial relevance to generative art: simple rules can produce complex, self-similar visual worlds. The spread of fractal imagery during the 1980s and 1990s also helps connect mathematical computation and visual invention in public culture.
Read original source · Les objets fractals (1975); The Fractal Geometry of Nature (1982) ↗
Research master · p. 10 · #027 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗1979Ars Electronica Festival
+
Ars Electronica Festival
Hannes Leopoldseder; Hubert Bognermayr; Herbert W. Franke; others
Beginning in 1979, Ars Electronica places art, technology and social questions on a shared public platform. Its continuing presentation of artificial life, robotics, data art and AI art provides long-term cultural infrastructure for practices that otherwise might remain isolated technological experiments.
Read original source · Ars Electronica Festival ↗
Research master · p. 11 · #028 · Research claims await independent verification
Open in the interactive timeline ↗1986Biomorph evolutionary program
+
Biomorph evolutionary program
Richard Dawkins
Dawkins's Biomorph program uses recursive rules and human selection to demonstrate cumulative evolution. It is not machine-learning art, but belongs alongside the evolutionary practices of Sims and Latham in a history of interactive evolutionary aesthetics, where creators select descendants within a generative space instead of drawing each result directly.
Read original source · Biomorph evolutionary program ↗
Research master · p. 11 · #029 · Research claims await independent verification
Open in the interactive timeline ↗1986Learning representations by back-propagating errors
+
Learning representations by back-propagating errors
Rumelhart, Hinton & Williams
Backpropagation allows hidden layers to learn internal representations useful for a task. It is not itself a technique for visual generation, but becomes one of the mechanisms underlying later deep-learning vision systems. Features need no longer be entirely designed by hand; they can form through training.
Read original source · Learning representations by back-propagating errors ↗
Research master · p. 11 · #030 · Research claims await independent verification
Open in the interactive timeline ↗1987Boids
+
Boids
Craig Reynolds
Boids simulates collective motion using local rules such as separation, alignment and cohesion. Its influence extends across animation, artificial life and generative art. The system shows how visual complexity can emerge from distributed rules rather than from a central designer specifying the behavior of the whole.
Read original source · Boids ↗
Research master · p. 12 · #031 · Research claims await independent verification
Open in the interactive timeline ↗1988–1992FormGrow / Mutator evolutionary systems
+
FormGrow / Mutator evolutionary systems
William Latham; Stephen Todd
Latham and Todd develop evolutionary systems for generating three-dimensional forms in settings including IBM UK. Their approach shifts artistic work from direct shaping toward defining generative rules and selecting mutations. The master treats this as a historical analogy for contemporary latent-space exploration and workflows built around variation.
Read original source · FormGrow / Mutator evolutionary systems ↗
Research master · p. 12 · #032 · Research claims await independent verification
Open in the interactive timeline ↗1991Artificial Evolution for Computer Graphics
+
Artificial Evolution for Computer Graphics
Karl Sims
Sims uses genetic algorithms and interactive selection to generate complex images, textures and animation. The machine produces candidates and users guide the next round through selection. This places aesthetic judgment inside the generative loop, rather than reserving it for evaluation after an artwork has been completed.
Read original source · Artificial Evolution for Computer Graphics ↗
Research master · p. 12 · #033 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗1994Evolving Virtual Creatures
+
Evolving Virtual Creatures
Karl Sims
Sims extends generation beyond still images by evolving the body structures and control networks of virtual creatures together. Generative practice encompasses dynamic behavior and artificial life, providing context for later work in games, simulation art and embodied generative systems.
Read original source · Evolving Virtual Creatures ↗
Research master · p. 12 · #034 · Research claims await independent verification
Open in the interactive timeline ↗1995 / 2000Pyramid-Based Texture Analysis/Synthesis; Parametric Texture Model
+
Pyramid-Based Texture Analysis/Synthesis; Parametric Texture Model
David Heeger; James Bergen; Javier Portilla; Eero Simoncelli
These texture-synthesis methods reconstruct appearance through multiscale filter responses and statistical constraints. Gatys's later neural style transfer uses statistics of deep-network features to represent style. The earlier tradition of statistical texture analysis is therefore an important technical antecedent, without implying that aesthetic style is exhausted by such statistics.
Read original source · Pyramid-Based Texture Analysis/Synthesis; Parametric Texture Model ↗
Research master · p. 13 · #035 · Research claims await independent verification
Open in the interactive timeline ↗1998Gradient-Based Learning Applied to Document Recognition
+
Gradient-Based Learning Applied to Document Recognition
LeCun et al.
LeNet-5 and related work integrate convolution, shared weights and gradient-based training into a stable visual architecture. Later developments including AlexNet, convolutional encoders in generative models and visual feature extraction inherit elements of this tradition. The node documents technical conditions for subsequent visual learning.
Read original source · Gradient-Based Learning Applied to Document Recognition ↗
Research master · p. 13 · #036 · Research claims await independent verification
Open in the interactive timeline ↗2001Image Quilting for Texture Synthesis and Transfer
+
Image Quilting for Texture Synthesis and Transfer
Alexei A. Efros; William T. Freeman
Image Quilting selects and joins patches from sample textures to synthesize larger textures, and also demonstrates texture transfer. It is not deep learning. It documents an important earlier route for reconstructing visual appearance from existing image material, preceding the widespread use of neural generative models.
Read original source · Image Quilting for Texture Synthesis and Transfer ↗
Research master · p. 13 · #037 · Research claims await independent verification
Open in the interactive timeline ↗2001Processing
+
Processing
Casey Reas; Ben Fry
Processing makes drawing, interaction and program structure accessible within a creative-coding environment. Although not AI, it becomes important infrastructure for generative art in the 2000s. The master relates its culture of programmable visual practice to later use of notebooks, node workflows and prompt scripting.
Read original source · Processing ↗
Research master · p. 13 · #038 · Research claims await independent verification
Open in the interactive timeline ↗2006A Fast Learning Algorithm for Deep Belief Nets
+
A Fast Learning Algorithm for Deep Belief Nets
Hinton, Osindero & Teh
This work is an important node in the revival of deep learning during the 2000s. It helps renew research into learning multilayer representations and creates conditions for later large-scale visual deep learning. Its historical role concerns representation learning rather than the direct production of artworks.
Read original source · A Fast Learning Algorithm for Deep Belief Nets ↗
Research master · p. 14 · #039 · Research claims await independent verification
Open in the interactive timeline ↗2006–2012The Painting Fool
+
The Painting Fool
Simon Colton
The Painting Fool is a long-running computational-creativity project that emphasizes a system's creative decisions and accounts of its own process. It connects discussions of symbolic and computational creativity with the cultural questions later raised by data-driven generative models.
Read original source · The Painting Fool ↗
Research master · p. 14 · #040 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗2009ImageNet: A Large-Scale Hierarchical Image Database
+
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng; Wei Dong; Richard Socher; Li-Jia Li; Kai Li; Li Fei-Fei
ImageNet connects WordNet's semantic hierarchy with millions of images and uses crowdsourcing for annotation. The master interprets this as a Dataset Turn: cultural and everyday images become organized as scalable resources for machine learning and comparison, alongside their existence as things people view. This turn is an editorial framework, not a neutral period label.
Read original source · ImageNet: A Large-Scale Hierarchical Image Database ↗
Research master · p. 14 · #041 · Research claims await independent verification
Open in the interactive timeline ↗2010ILSVRC
+
ILSVRC
ImageNet team; international computer vision community
ILSVRC standardizes subsets of ImageNet into an annual visual-recognition competition. Data become not only training material, but a framework for measuring and comparing research performance. AlexNet's 2012 results are recognized and amplified within this institutional structure of datasets, benchmarks and competitive evaluation.
Read original source · ILSVRC ↗
Research master · p. 14 · #042 · Research claims await independent verification
Open in the interactive timeline ↗2012AlexNet
+
AlexNet
Krizhevsky, Sutskever & Hinton
AlexNet substantially outperforms established methods on ImageNet classification and helps make deep convolutional networks central to computer vision. The master reads this as a convergence of dataset scale, GPU computation and learned representations, extending the Dataset Turn into the production of new visual capabilities.
Read original source · AlexNet ↗
Research master · p. 15 · #043 · Research claims await independent verification
Open in the interactive timeline ↗2013Auto-Encoding Variational Bayes
+
Auto-Encoding Variational Bayes
Kingma & Welling
The variational autoencoder establishes a differentiable framework for generative latent-variable modeling. For this genealogy, a significant legacy is the movement of latent space from statistical terminology into visual culture: images can be compressed into representations, reconstructed, and newly sampled within a learned generative space.
Read original source · Auto-Encoding Variational Bayes ↗
Research master · p. 15 · #044 · Research claims await independent verification
Open in the interactive timeline ↗2014Generative Adversarial Nets
+
Generative Adversarial Nets
Goodfellow et al.
GANs organize generation through learning data distributions. In contrast with a rule-based system such as AARON, the logic producing an image is learned from training data rather than mainly being specified by an artist. The master identifies a change in how machine images are understood: outputs are treated as samples from a learned distribution.
Read original source · Generative Adversarial Nets ↗
Research master · p. 15 · #045 · Research claims await independent verification
Open in the interactive timeline ↗2014Conditional Generative Adversarial Nets
+
Conditional Generative Adversarial Nets
Mehdi Mirza; Simon Osindero
Conditional GANs supply additional information to both generator and discriminator, offering a general framework for controlled generation. Class-conditioned synthesis, text-to-image GANs and forms of structurally conditioned generation build on this approach. The node concerns the development of conditions that guide what a model generates.
Read original source · Conditional Generative Adversarial Nets ↗
Research master · p. 15 · #046 · Research claims await independent verification
Open in the interactive timeline ↗2015LAPGAN
+
LAPGAN
Emily Denton; Soumith Chintala; Arthur Szlam; Rob Fergus
LAPGAN generates image detail at multiple scales using a Laplacian pyramid. It is an important attempt at higher-resolution synthesis before Progressive GAN. The work illustrates a continuing problem in image generation: how to extend quality through the coordinated construction of images across layers or scales.
Read original source · LAPGAN ↗
Research master · p. 16 · #047 · Research claims await independent verification
Open in the interactive timeline ↗2015A Neural Algorithm of Artistic Style
+
A Neural Algorithm of Artistic Style
Gatys, Ecker & Bethge
Neural style transfer represents content through deep convolutional features and describes style through feature statistics, then synthesizes a new image. It does not establish that artistic style is fundamentally identical to those statistics. Its cultural importance lies in making the extraction, computation and transfer of style appear tangible within visual practice.
Read original source · A Neural Algorithm of Artistic Style ↗
Research master · p. 16 · #048 · Research claims await independent verification
Open in the interactive timeline ↗2015DeepDream
+
DeepDream
Alexander Mordvintsev; Christopher Olah; Mike Tyka / Google
DeepDream amplifies internal network activations to make recognized patterns visible. It is not contemporary text-to-image generation, but gives a broad public an early encounter with a recognizable visual aesthetic arising from neural representations. Artists, media and online communities quickly incorporate that aesthetic into their practices and discussions.
Read original source · DeepDream ↗
Research master · p. 16 · #049 · Research claims await independent verification
Open in the interactive timeline ↗2015DRAW
+
DRAW
Karol Gregor; Ivo Danihelka; Alex Graves; Danilo Rezende; Daan Wierstra
DRAW combines a variational autoencoder, recurrent networks and differentiable attention, generating images through successive operations that read from and write to a canvas. Although it does not become the dominant architecture of later products, it demonstrates an early combination of attention and iterative image generation.
Research master · p. 16 · #050 · Research claims await independent verification
Open in the interactive timeline ↗2015DCGAN
+
DCGAN
Radford, Metz & Chintala
DCGAN establishes architectural practices adopted by many later visual GANs and demonstrates that directions in latent space can correspond to semantic changes. It brings manipulation of latent representations closer to a practical creative workflow, in which users explore meaningful variations instead of selecting only among unrelated generated images.
Read original source · DCGAN ↗
Research master · p. 17 · #051 · Research claims await independent verification
Open in the interactive timeline ↗2016pix2pix
+
pix2pix
Isola et al.
pix2pix learns from paired images to generate a target image from inputs such as label maps, edge maps and sketches. It extends generation beyond sampling from noise toward transformation conditioned on visual structure. This provides a clear technical lineage for later editing, redrawing and structural control.
Read original source · pix2pix ↗
Research master · p. 17 · #052 · Research claims await independent verification
Open in the interactive timeline ↗2016Prisma mobile application
+
Prisma mobile application
Prisma Labs
Applications such as Prisma turn neural style transfer into a visual interface accessible to ordinary users. The historical point here is not a new algorithm, but a change in use: style becomes something that can be invoked as an everyday visual effect through a consumer application.
Read original source · Prisma mobile application ↗
Research master · p. 17 · #053 · Research claims await independent verification
Open in the interactive timeline ↗2016–2017StackGAN: Text to Photo-realistic Image Synthesis with Stacked GANs
+
StackGAN: Text to Photo-realistic Image Synthesis with Stacked GANs
Han Zhang; Tao Xu; Hongsheng Li; Shaoting Zhang; Xiaogang Wang; Xiaolei Huang; Dimitris Metaxas
StackGAN uses a two-stage process that first turns a description into low-resolution structure and then refines it into a higher-resolution image. It is a significant branch of text-to-image research before DALL·E and Imagen, demonstrating that prompt-driven image generation has a technical history preceding diffusion models.
Read original source · StackGAN: Text to Photo-realistic Image Synthesis with Stacked GANs ↗
Research master · p. 17 · #054 · Research claims await independent verification
Open in the interactive timeline ↗2017Progressive Growing of GANs
+
Progressive Growing of GANs
Karras et al.
Progressive GAN increases resolution layer by layer to improve training stability and image detail. Its high-resolution synthetic portraits become culturally striking, and the work leads into StyleGAN. The node connects an architectural strategy for scaling generation with the changing public visibility of synthetic faces.
Read original source · Progressive Growing of GANs ↗
Research master · p. 18 · #055 · Research claims await independent verification
Open in the interactive timeline ↗2017Attention Is All You Need
+
Attention Is All You Need
Vaswani et al.
The Transformer is initially a sequence-modeling architecture, not a paper on image generation. Later systems including CLIP, DALL·E, DiT, MMDiT and native multimodal models nevertheless draw directly or indirectly on attention and Transformer approaches. Its inclusion documents a cross-domain technical condition rather than an artwork or image model in itself.
Read original source · Attention Is All You Need ↗
Research master · p. 18 · #056 · Research claims await independent verification
Open in the interactive timeline ↗2017CycleGAN
+
CycleGAN
Zhu et al.
CycleGAN uses cycle consistency to learn translation between unpaired image domains, extending tasks such as changes of season, transformations between painting and photographic domains, and changes in object appearance. It helps develop the idea that a visual domain or style can be learned statistically and transformed into another.
Read original source · CycleGAN ↗
Research master · p. 18 · #057 · Research claims await independent verification
Open in the interactive timeline ↗2017Neural Discrete Representation Learning
+
Neural Discrete Representation Learning
Aaron van den Oord; Oriol Vinyals; Koray Kavukcuoglu
VQ-VAE learns a discrete codebook that compresses high-dimensional inputs into combinable representations. DALL·E, VQGAN and several autoregressive image models subsequently draw on the idea of visual tokens. The node concerns the representation of images in units that can participate in learned sequences and generative processes.
Read original source · Neural Discrete Representation Learning ↗
Research master · p. 18 · #058 · Research claims await independent verification
Open in the interactive timeline ↗2017–2018AttnGAN: Fine-Grained Text to Image Generation with Attentional GANs
+
AttnGAN: Fine-Grained Text to Image Generation with Attentional GANs
Tao Xu; Pengchuan Zhang; Qiuyuan Huang; Han Zhang; Zhe Gan; Xiaolei Huang; Xiaodong He
AttnGAN aligns words with visual regions through attention and a deep attentional multimodal similarity model, improving details corresponding to complex descriptions. It provides an important pre-diffusion example of the problems now discussed as adherence to prompts and alignment between textual and visual information.
Read original source · AttnGAN: Fine-Grained Text to Image Generation with Attentional GANs ↗
Research master · p. 19 · #059 · Research claims await independent verification
Open in the interactive timeline ↗2018BigGAN
+
BigGAN
Brock, Donahue & Simonyan
BigGAN scales GAN training for class-conditioned ImageNet generation and improves image fidelity. It reinforces an observation that recurs in the foundation-model period: model size, data and computational scale can themselves contribute to generative quality. The historical argument concerns this scaling relationship, not an automatic equation of scale with artistic value.
Read original source · BigGAN ↗
Research master · p. 19 · #060 · Research claims await independent verification
Open in the interactive timeline ↗2018StyleGAN
+
StyleGAN
Karras, Laine & Aila
StyleGAN introduces a mapping network and a style-based generator that offer more intuitive control over visual attributes at different scales. It broadens the cultural influence of navigable latent spaces, supporting face synthesis, latent editing and AI portrait practices. The model's technical use of style should remain distinct from the full art-historical concept.
Read original source · StyleGAN ↗
Research master · p. 19 · #061 · Research claims await independent verification
Open in the interactive timeline ↗2018–2019Ganbreeder / Artbreeder
+
Ganbreeder / Artbreeder
Joel Simon; Artbreeder community
Joel Simon's Ganbreeder develops into Artbreeder, turning GAN latent spaces into interfaces for browsing, mixing, inheriting and sharing visual forms. Before prompt culture becomes dominant, this is a significant example of latent culture, where interaction with generative models is organized around visual variation and social circulation.
Read original source · Ganbreeder / Artbreeder ↗
Research master · p. 19 · #062 · Research claims await independent verification
Open in the interactive timeline ↗2018-10Portrait of Edmond de Belamy
+
Portrait of Edmond de Belamy
Obvious (Hugo Caselles-Dupré; Pierre Fautrel; Gauthier Vernier); Robbie Barrat lineage controversy
Christie's sale of Portrait of Edmond de Belamy for $432,500 makes GAN art an international media subject. It also brings disputes over the provenance of code, authorship and the attribution of technical labor into view. The artwork's value chain includes models, software, an artistic collective and the institutions of the art market.
Read original source · Portrait of Edmond de Belamy ↗
Research master · p. 20 · #063 · Research claims await independent verification
Open in the interactive timeline ↗2019AI: More than Human
+
AI: More than Human
Barbican curatorial team; international artists/researchers
The Barbican exhibition places historical automata, machine learning, robotics and contemporary AI art within a shared narrative space. It shows AI art becoming an established subject of institutional exhibition around 2019, providing context for the controversies that follow the wider spread of generative image products in 2022.
Read original source · AI: More than Human ↗
Research master · p. 20 · #064 · Research claims await independent verification
Open in the interactive timeline ↗2019Memories of Passersby I
+
Memories of Passersby I
Mario Klingemann
Klingemann presents and sells an installation that continuously generates images. The artwork is therefore not simply one AI-produced picture, but a working generative system and its potentially continuing output. This raises questions about the status of the art object, collecting and authorship beyond those posed by a single generated image.
Read original source · Memories of Passersby I ↗
Research master · p. 20 · #065 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗2019Generating Diverse High-Fidelity Images with VQ-VAE-2
+
Generating Diverse High-Fidelity Images with VQ-VAE-2
Ali Razavi; Aaron van den Oord; Oriol Vinyals
VQ-VAE-2 combines hierarchical discrete latent variables with autoregressive priors to produce high-quality images for its period. It supplies technical background for approaches such as DALL·E, in which visual representations are tokenized and modeled within larger generative systems.
Read original source · Generating Diverse High-Fidelity Images with VQ-VAE-2 ↗
Research master · p. 20 · #066 · Research claims await independent verification
Open in the interactive timeline ↗2019SPADE paper / GauGAN demo
+
SPADE paper / GauGAN demo
Taesung Park; Ming-Yu Liu; Ting-Chun Wang; Jun-Yan Zhu / NVIDIA
SPADE converts semantic layouts into high-quality images, while NVIDIA's GauGAN demonstration presents its potential as an interactive creative tool. It anticipates later interest in layout control, ControlNet and design-oriented generation, where a user specifies spatial organization rather than relying on text alone.
Read original source · SPADE paper / GauGAN demo ↗
Research master · p. 21 · #067 · Research claims await independent verification
Open in the interactive timeline ↗2019Analyzing and Improving the Image Quality of StyleGAN
+
Analyzing and Improving the Image Quality of StyleGAN
Tero Karras; Samuli Laine; Miika Aittala; Janne Hellsten; Jaakko Lehtinen; Timo Aila
StyleGAN2 addresses visual artifacts in StyleGAN and improves the mapping from latent representations to images. It strengthens techniques for inversion, editing and attribute manipulation, helping develop a workflow in which an image can be generated, mapped back into a model's representation and edited within that space.
Read original source · Analyzing and Improving the Image Quality of StyleGAN ↗
Research master · p. 21 · #068 · Research claims await independent verification
Open in the interactive timeline ↗2020Denoising Diffusion Probabilistic Models
+
Denoising Diffusion Probabilistic Models
Ho, Jain & Abbeel
DDPM demonstrates high-quality image synthesis with diffusion probabilistic models. In contrast with a GAN's direct generation, diffusion constructs an image through an iterative process of reversing noise. This approach subsequently becomes a major technical basis for text-to-image systems after 2022.
Read original source · Denoising Diffusion Probabilistic Models ↗
Research master · p. 21 · #069 · Research claims await independent verification
Open in the interactive timeline ↗2020–2021Taming Transformers / VQGAN
+
Taming Transformers / VQGAN
Esser, Rombach & Ommer
VQGAN combines convolutional encoding of local visual structure with a Transformer's capacity for broader modeling, producing high-quality discrete visual representations. It directly influences early VQGAN+CLIP art practices and contributes technical experience to later generative models operating within compressed latent spaces.
Read original source · Taming Transformers / VQGAN ↗
Research master · p. 21 · #070 · Research claims await independent verification
Open in the interactive timeline ↗2020-04Jukebox: A Generative Model for Music
+
Jukebox: A Generative Model for Music
Prafulla Dhariwal et al.; OpenAI
Jukebox uses hierarchical VQ-VAEs and autoregressive priors to generate music and approximate vocals. It is distinct from later text-driven music platforms. Its relevance is cross-media: learning generative capacities from large collections of cultural recordings is already becoming a paradigm that extends beyond images.
Read original source · Jukebox: A Generative Model for Music ↗
Research master · p. 22 · #071 · Research claims await independent verification
Open in the interactive timeline ↗2020-05Language Models are Few-Shot Learners
+
Language Models are Few-Shot Learners
Tom B. Brown et al.; OpenAI
GPT-3 is not an image model, but helps elevate the prompt from ordinary input text into an interface for specifying many different tasks. The master connects later systems such as DALL·E 3, conversational image generation and language-driven editing to this broader change: natural language becomes a control layer for models.
Read original source · Language Models are Few-Shot Learners ↗
Research master · p. 22 · #072 · Research claims await independent verification
Open in the interactive timeline ↗2021Botto
+
Botto
Mario Klingemann; Botto community
Botto links the generation of candidate images with community voting and auctions in a continuing creative loop. It combines generative models, collective taste, token-based governance and the art market. This makes it a case for studying machine authorship alongside the organization of cultural production through platforms and communities.
Read original source · Botto ↗
Research master · p. 22 · #073 · Research claims await independent verification
Open in the interactive timeline ↗2021VQGAN+CLIP community workflow
+
VQGAN+CLIP community workflow
Community practitioners; Katherine Crowson and open-source notebook ecosystem
Practitioners combine VQGAN's image synthesis with CLIP's text–image similarity to guide generation using language. This is a distributed practical arrangement rather than one canonical paper. Prompt recipes, artists' names, seeds, notebooks and iteration settings become part of a shared creative vocabulary within the open-source notebook ecosystem.
Read original source · VQGAN+CLIP community workflow ↗
Research master · p. 22 · #074 · Research claims await independent verification
Open in the interactive timeline ↗2021-02Zero-Shot Text-to-Image Generation
+
Zero-Shot Text-to-Image Generation
Aditya Ramesh; Mikhail Pavlov; Gabriel Goh; Scott Gray; Chelsea Voss et al. / OpenAI
DALL·E uses an autoregressive Transformer to model text and image tokens together, demonstrating broad compositional capabilities without task-specific examples. It makes the idea of producing a new image from a sentence publicly legible as a generative paradigm, connecting research in token modeling with a new interface to images.
Read original source · Zero-Shot Text-to-Image Generation ↗
Research master · p. 23 · #075 · Research claims await independent verification
Open in the interactive timeline ↗2021-02Learning Transferable Visual Models From Natural Language Supervision
+
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford et al. / OpenAI
CLIP learns image–text alignment from internet image–text pairs, allowing language to indicate visual concepts in a zero-shot setting. It becomes a key component of early prompt-driven art practices including VQGAN+CLIP and Disco Diffusion, while bringing web-scale image–text data into the center of generative culture.
Read original source · Learning Transferable Visual Models From Natural Language Supervision ↗
Research master · p. 23 · #076 · Research claims await independent verification
Open in the interactive timeline ↗2021-05Diffusion Models Beat GANs
+
Diffusion Models Beat GANs
Dhariwal & Nichol
Architectural improvements and classifier guidance allow diffusion models to outperform prominent GAN approaches on ImageNet synthesis. The master identifies this as a technical turning point in the transition from GAN-centered image generation toward diffusion. The comparison concerns a particular research benchmark and period, rather than every possible artistic or technical application.
Read original source · Diffusion Models Beat GANs ↗
Research master · p. 23 · #077 · Research claims await independent verification
Open in the interactive timeline ↗2021-10-29Disco Diffusion notebook/system
+
Disco Diffusion notebook/system
Somnai; Gandamu; Disco Diffusion contributors/community
From October 2021, Disco Diffusion develops a reusable Colab workflow combining diffusion, CLIP guidance, cutouts, image prompts and animation parameters. Its art-historical importance lies as much in practice as in foundational research: large numbers of users adopt prompt engineering and iterative parameter choices as forms of creative work.
Read original source · Disco Diffusion notebook/system ↗
Research master · p. 23 · #078 · Research claims await independent verification
Open in the interactive timeline ↗2021-12GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
+
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Alex Nichol; Prafulla Dhariwal et al. / OpenAI
GLIDE compares CLIP guidance with classifier-free guidance and demonstrates photorealistic text-to-image generation and inpainting. It is a precursor to DALL·E 2 and later diffusion-based products. The combination of generation and editing is especially relevant to the development of interfaces for modifying existing visual material.
Research master · p. 24 · #079 · Research claims await independent verification
Open in the interactive timeline ↗2021-12 / 2022-06Latent Diffusion Models
+
Latent Diffusion Models
Rombach et al.
Latent diffusion operates in the compressed space of a pretrained autoencoder and accepts conditions such as text through cross-attention. It reduces computational costs and provides the direct technical foundation for Stable Diffusion, helping make an open text-to-image ecosystem practicable on consumer graphics hardware.
Read original source · Latent Diffusion Models ↗
Research master · p. 24 · #080 · Research claims await independent verification
Open in the interactive timeline ↗2022LAION-5B: An open large-scale dataset for training next generation image-text models
+
LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann et al. / LAION
LAION-5B makes a large image–text pair index and filtering methods available, providing data infrastructure relevant to Stable Diffusion and related ecosystems. It also becomes central to disputes about training sources, artists' names and copyright. The master reads this as a shift in which the dataset itself becomes a political object; an index does not confer ownership of linked images.
Research master · p. 24 · #081 · Research claims await independent verification
Open in the interactive timeline ↗2022Hierarchical Text-Conditional Image Generation with CLIP Latents / DALL·E 2 product
+
Hierarchical Text-Conditional Image Generation with CLIP Latents / DALL·E 2 product
Aditya Ramesh; Prafulla Dhariwal; Alex Nichol; Casey Chu / OpenAI
DALL·E 2 combines CLIP representations with a diffusion-based image prior and decoder, improving realism and correspondence with text. Its research preview, beta and later access without a waitlist expand during 2022. The master emphasizes this product transition as a step in the international popularization of text-to-image generation.
The master gives 2022-03. Product and paper release months require reconciliation; this view uses year precision.
Research master · p. 25 · #082 · Research claims await independent verification
Open in the interactive timeline ↗2022–2023Civitai model-sharing platform
+
Civitai model-sharing platform
Civitai founders/team and model-sharing community
Platforms such as Civitai turn fine-tuned models and configurations from the Stable Diffusion ecosystem into circulating cultural objects. Generative capability diversifies from a single foundation model into a large community market. This also complicates questions of provenance, style imitation, models of identifiable people and consent.
Read original source · Civitai model-sharing platform ↗
Research master · p. 27 · #091 · Research claims await independent verification
Open in the interactive timeline ↗2022–2023LoRA: Low-Rank Adaptation of Large Language Models + diffusion community adoption
+
LoRA: Low-Rank Adaptation of Large Language Models + diffusion community adoption
Edward Hu et al. (LoRA); Stable Diffusion open-source community
The original LoRA paper appears in 2021, reducing fine-tuning costs by freezing a main model and training low-rank update matrices. During 2022–2023, the Stable Diffusion community adapts this into small exchangeable modules for styles, characters, clothing, poses and likenesses. Style and identity can circulate as lightweight model files as well as names in prompts.
Research master · p. 27 · #092 · Research claims await independent verification
Open in the interactive timeline ↗2022-05Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
+
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia et al. / Google Research
Imagen uses a large text encoder and diffusion decoders. The research highlights gains in image–text alignment from scaling the language model. It develops a direction that later multimodal systems also pursue: improvements in understanding language can change the visual capabilities and controllability of image generation.
Research master · p. 25 · #083 · Research claims await independent verification
Open in the interactive timeline ↗2022-06Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
+
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Jiahui Yu et al. / Google Research
Parti explores an autoregressive approach alongside diffusion, generating image tokens as another kind of language. It shows that the expansion of text-to-image generation in 2022 did not follow a single technical route. Token-based modeling remains a parallel way to organize the relation between descriptions and visual outputs.
Read original source · Scaling Autoregressive Models for Content-Rich Text-to-Image Generation ↗
Research master · p. 25 · #084 · Research claims await independent verification
Open in the interactive timeline ↗2022-06Diffusers library
+
Diffusers library
Hugging Face Diffusers team
Diffusers organizes schedulers, pipelines, model components and training scripts as reusable software modules, lowering barriers to deploying and extending diffusion models. It is an intermediate layer through which research models become ecosystems of applications, making software infrastructure an important part of the genealogy of generative practice.
Read original source · Diffusers library ↗
Research master · p. 25 · #085 · Research claims await independent verification
Open in the interactive timeline ↗2022-07Midjourney open beta / Discord platform
+
Midjourney open beta / Discord platform
David Holz; Midjourney team/community
Midjourney's historical importance extends beyond model quality. Its Discord interface lets participants see one another's prompts and results, supporting social learning, the circulation of styles and prompt imitation. Text-to-image generation becomes a public visual culture as well as an individual creative tool.
Read original source · Midjourney open beta / Discord platform ↗
Research master · p. 26 · #086 · Research claims await independent verification
Open in the interactive timeline ↗2022-08–10stable-diffusion-webui
+
stable-diffusion-webui
AUTOMATIC1111 and open-source contributors
The WebUI brings text-to-image generation, image-to-image transformation, inpainting, samplers and extensions into one interface, later incorporating LoRA and ControlNet. Community software of this kind is essential to the practical transformation of downloadable model weights into tools that large numbers of creators can actually use.
Read original source · stable-diffusion-webui ↗
Research master · p. 26 · #087 · Research claims await independent verification
Open in the interactive timeline ↗2022-08-02An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
+
An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Rinon Gal; Yuval Alaluf; Yuval Atzmon; Or Patashnik; Amit Bermano; Gal Chechik; Daniel Cohen-Or
Textual Inversion learns a new text embedding without retraining an entire model. Visual concepts, objects and styles can be represented through a new composable word within the system. This is an effective method of personalization, not a claim that the nature of artistic style can be reduced to a token.
Research master · p. 28 · #093 · Research claims await independent verification
Open in the interactive timeline ↗2022-08-22Stable Diffusion public model release
+
Stable Diffusion public model release
Robin Rombach; Patrick Esser; CompVis; Stability AI; Runway collaborators
Stable Diffusion builds on latent diffusion and makes model weights publicly available. The master emphasizes a change in access: users can download, fine-tune, fork, combine and run models locally, rather than only access a corporate API. Community systems including LoRA workflows, ControlNet, Civitai and ComfyUI develop within this broader ecosystem.
Read original source · Stable Diffusion public model release ↗
Research master · p. 26 · #088 · Research claims await independent verification
Open in the interactive timeline ↗2022-08-25DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
+
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
Nataniel Ruiz; Yuanzhen Li; Varun Jampani; Yael Pritch; Michael Rubinstein; Kfir Aberman
DreamBooth fine-tunes a pretrained text-to-image model using a small set of subject images, associating a unique identifier with that subject. It brings preservation of identity into the generative process and supplies technical context for later customization of people, characters and products, as well as disputes concerning likeness.
Research master · p. 28 · #094 · Research claims await independent verification
Open in the interactive timeline ↗2022-09Théâtre D’opéra Spatial (Space Opera Theater)
+
Théâtre D’opéra Spatial
Jason M. Allen; Midjourney
Jason Allen enters and wins a competition with a work made using Midjourney and subsequent editing. The event becomes a cultural marker of what the master calls the 2022 Political Turn: discussion shifts from technical demonstrations toward creative labor, disclosure rules, authorship and fairness in artistic competition.
Read original source · Théâtre D’opéra Spatial ↗
This official 2023 registration review recounts the work and its 2022 award.
Research master · p. 26 · #089 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗2022-11Unsupervised
+
Unsupervised
Refik Anadol Studio; MoMA
Unsupervised uses data from MoMA's collection to train or drive a generative visual system presented on a large real-time screen. It makes relations among datasets, museum collections and generative models visible as artistic material. The institutional collection participates in both the system's input and the public interpretation of its output.
Read original source · Unsupervised ↗
Research master · p. 27 · #090 · Research claims await independent verification
Open in the interactive timeline ↗2022-11 / 2023InstructPix2Pix: Learning to Follow Image Editing Instructions
+
InstructPix2Pix: Learning to Follow Image Editing Instructions
Tim Brooks; Aleksander Holynski; Alexei A. Efros
InstructPix2Pix uses GPT-3 and Stable Diffusion to synthesize training examples, then trains a diffusion model conditioned on editing instructions. It shifts the interaction from making a picture in a single generation toward changing an existing image through language, connecting generative modeling with iterative visual editing.
Read original source · InstructPix2Pix: Learning to Follow Image Editing Instructions ↗
Research master · p. 28 · #095 · Research claims await independent verification
Open in the interactive timeline ↗2023ComfyUI
+
ComfyUI
ComfyAnonymous and open-source contributors
ComfyUI makes a Stable Diffusion workflow explicit as a graph. Users arrange models and control modules within a visual system rather than only entering a prompt. The master interprets this as a shift from using a model toward designing a production pipeline, where organization and reuse of processes become part of image-making.
Read original source · ComfyUI ↗
Research master · p. 29 · #096 · Research claims await independent verification
Open in the interactive timeline ↗2023AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
+
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Yuwei Guo et al.
AnimateDiff introduces motion modeling as a pluggable component within personalized text-to-image diffusion systems. Existing character models, styles and community assets can participate in animated outputs. This extends generation from still images toward video while retaining parts of the earlier image-model ecosystem.
Research master · p. 31 · #106 · Research claims await independent verification
Open in the interactive timeline ↗2023-02Adding Conditional Control to Text-to-Image Diffusion Models
+
Adding Conditional Control to Text-to-Image Diffusion Models
Lvmin Zhang; Anyi Rao; Maneesh Agrawala
ControlNet freezes a pretrained diffusion backbone and learns spatial conditions through an additional network. It extends control from textual semantics toward geometry and composition. This is one technical condition for more specialized production workflows, in which artists or designers need to specify structure as well as subject matter.
Read original source · Adding Conditional Control to Text-to-Image Diffusion Models ↗
Research master · p. 29 · #097 · Research claims await independent verification
Open in the interactive timeline ↗2023-02T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
+
T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
Chong Mou; Xintao Wang; Liangbin Xie; Yanze Wu; Jian Zhang; Zhongang Qi; Ying Shan; Xiaohu Qie
Alongside ControlNet, T2I-Adapter exemplifies a move toward fine-grained external control through modular attachments without fully retraining a foundation model. Image production becomes a combination of a base model and adapters. The node situates controllability in a growing ecosystem of interoperable components.
Research master · p. 29 · #098 · Research claims await independent verification
Open in the interactive timeline ↗2023-03-21Adobe Firefly / Generative Fill
+
Adobe Firefly / Generative Fill
Adobe Firefly / Adobe Research and Creative Cloud teams
Firefly places generative AI inside established creative-software infrastructure. Its launch materials connect the product with questions about training sources, creator compensation, training preferences and content credentials. The master treats this as a meeting between model development and the governance of professional creative production; product commitments must be distinguished from independently demonstrated outcomes.
Read original source · Adobe Firefly / Generative Fill ↗
Research master · p. 29 · #099 · Research claims await independent verification
Open in the interactive timeline ↗2023-05Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
+
Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
Xingang Pan et al.; Max Planck Institute for Informatics / Saarland research team
DragGAN uses a GAN's feature space for point-based manipulation, demonstrating direct interaction with a generative image manifold. Although diffusion becomes increasingly prominent, this approach remains relevant to the development of editing interfaces in which users move elements visually instead of specifying every change through text.
Research master · p. 30 · #100 · Research claims await independent verification
Open in the interactive timeline ↗2023-06Gen-2 multimodal video generation
+
Gen-2 multimodal video generation
Runway
Runway Gen-2 supports short-video generation conditioned on text and images. It is an important example of video generation becoming a product in 2023. Creative concerns expand from a single composition toward motion, camera behavior, temporal consistency and the demands of audiovisual production.
Read original source · Gen-2 multimodal video generation ↗
Research master · p. 31 · #107 · Research claims await independent verification
Open in the interactive timeline ↗2023-07-26SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
+
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Robin Rombach; Stability AI research team and collaborators
SDXL develops composition, detail, style and resolution while retaining a model ecosystem that can be fine-tuned and combined with other tools. The master positions it within the Stable Diffusion community's movement from rapid experimentation toward more established production practices.
Read original source · SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis ↗
Research master · p. 30 · #101 · Research claims await independent verification
Open in the interactive timeline ↗2023-08IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
+
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Hu Ye; Jun Zhang; Sibo Liu; Xiao Han; Wei Yang / Tencent AI Lab
IP-Adapter uses decoupled cross-attention to add lightweight image-prompt capabilities to pretrained diffusion models. Reference-image control can coexist with text prompts and ControlNet, supporting workflows for portraits, fashion and identity that combine several types of guidance rather than relying on a description alone.
Research master · p. 30 · #102 · Research claims await independent verification
Open in the interactive timeline ↗2023-08Ideogram text-to-image model
+
Ideogram text-to-image model
Ideogram team
Ideogram attracts attention for rendering text within generated images. The node illustrates how text-to-image competition extends beyond visually appealing scenes toward typography, layouts, logos and commercial graphics. Design usability becomes a distinct area of model development and product differentiation.
Read original source · Ideogram text-to-image model ↗
Research master · p. 30 · #103 · Research claims await independent verification
Open in the interactive timeline ↗2023-09PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
+
PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Junsong Chen et al.
PixArt-α focuses on diffusion Transformers, pretrained text encoders and efficient training. It is an open research node leading into the wider development of Transformer- and flow-based image systems in 2024. Its significance concerns both architectural choices and the resources needed to train visual models.
Research master · p. 34 · #118 · Research claims await independent verification
Open in the interactive timeline ↗2023-09 / 2023-10DALL·E 3 / ChatGPT integration
+
DALL·E 3 / ChatGPT integration
OpenAI image generation team; ChatGPT product team
Integration with ChatGPT lets users discuss an idea and request changes while a language model organizes image prompts. The interface shifts from a prompt box toward conversation, reducing the need to master a particular model's prompting conventions. The genealogy follows this change in interaction as well as changes in image quality.
Read original source · DALL·E 3 / ChatGPT integration ↗
Research master · p. 37 · #128 · Research claims await independent verification
Open in the interactive timeline ↗2023-10Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
+
Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
Simian Luo et al.
Latent Consistency Models use consistency distillation in latent space to generate high-resolution images in a small number of steps, including two to four in the reported approach. Speed changes the possible interface: generation closer to real time can participate more readily in drawing, design, live presentation and interactive software.
Research master · p. 31 · #104 · Research claims await independent verification
Open in the interactive timeline ↗2023-11Pika 1.0
+
Pika 1.0
Pika Labs
Pika packages text-to-video, image-to-video and video modification as accessible creative tools. Alongside Runway, it helps move generative video from research and specialist experimentation toward consumer products. Its historical role includes the interface and availability of the tool, as well as the underlying model.
Read original source · Pika 1.0 ↗
Research master · p. 32 · #108 · Research claims await independent verification
Open in the interactive timeline ↗2023-11Stable Video Diffusion
+
Stable Video Diffusion
Stability AI / research collaborators
Stable Video Diffusion makes latent video diffusion available to researchers and developers, extending an open generative ecosystem from still images toward video. It also points toward increasingly shared infrastructure between image and video models, including software for arranging and adapting generation processes.
Read original source · Stable Video Diffusion ↗
Research master · p. 32 · #109 · Research claims await independent verification
Open in the interactive timeline ↗2023-12Midjourney V6
+
Midjourney V6
Midjourney team
Midjourney V6 develops prompt accuracy, coherence, image prompting and remix capabilities. It represents continuing iteration within a closed platform, where visual quality, aesthetic character and community interfaces evolve together. The node documents one product lineage alongside the open-weight and research-model histories elsewhere in the timeline.
Read original source · Midjourney V6 ↗
Research master · p. 31 · #105 · Research claims await independent verification
Open in the interactive timeline ↗2024Recraft V3
+
Recraft V3
Recraft team
Design-oriented systems such as Recraft extend generation from visual imagination toward brand assets, typography, vectors and deliverable design material. This differentiation shows image generation becoming a set of tools for distinct creative industries, rather than remaining one undifferentiated technical category.
Read original source · Recraft V3 ↗
Research master · p. 34 · #119 · Research claims await independent verification
Open in the interactive timeline ↗2024-01InstantID: Zero-shot Identity-Preserving Generation in Seconds
+
InstantID: Zero-shot Identity-Preserving Generation in Seconds
Qixun Wang et al.
InstantID combines facial embeddings with structural conditions to support efficient identity preservation in diffusion models. It marks a move from training specifically on a person, as in DreamBooth-style workflows, toward generation guided by references. The technique is directly relevant to later questions of likeness and consent.
Read original source · InstantID: Zero-shot Identity-Preserving Generation in Seconds ↗
Research master · p. 35 · #120 · Research claims await independent verification
Open in the interactive timeline ↗2024-02Sora
+
Sora
OpenAI Sora research and product teams
Sora shifts public expectations of generative models toward sustained scenes, camera motion and physical consistency. The master sees this as widening research on generative visual culture to encompass video models and film production. Expectations or claims about world simulation should be distinguished from independently demonstrated capabilities.
Research master · p. 32 · #110 · Research claims await independent verification
Open in the interactive timeline ↗2024-02Stable Cascade
+
Stable Cascade
Stability AI; Würstchen research lineage
Stable Cascade builds on the Würstchen approach, combining strong semantic compression with staged generation. Although it does not become the dominant community ecosystem, it documents exploration of efficiency, compressed representations and modular generation in the period following SDXL.
Read original source · Stable Cascade ↗
Research master · p. 35 · #121 · Research claims await independent verification
Open in the interactive timeline ↗2024-02Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
+
Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Stability AI research team
Stable Diffusion 3 combines a Transformer-based architecture with flow matching and emphasizes multi-subject prompts, text and image quality. It represents the further incorporation of Transformer approaches into mainstream image generation. The associated research paper provides a technical account distinct from the product announcement.
Research master · p. 35 · #122 · Research claims await independent verification
Open in the interactive timeline ↗2024-05Veo video generation model
+
Veo video generation model
Google DeepMind
Announced at Google I/O 2024, Veo is presented with capabilities including 1080p output, sequences longer than a minute and interpretation of cinematic language. The node records a shift in product ambitions from brief demonstrations toward questions of duration, camera work and control associated with audiovisual production.
Read original source · Veo video generation model ↗
Research master · p. 32 · #111 · Research claims await independent verification
Open in the interactive timeline ↗2024-05 / 2024-08Imagen 3
+
Imagen 3
Google DeepMind / Imagen team
Imagen 3 is introduced at Google I/O and subsequently enters Gemini and Vertex AI. The master treats this as an example of general-purpose AI platforms incorporating image generation as a standard capability within their model stacks, rather than maintaining it solely as a separate experimental product.
Read original source · Imagen 3 ↗
Research master · p. 35 · #123 · Research claims await independent verification
Open in the interactive timeline ↗2024-06Kling video generation model
+
Kling video generation model
Kuaishou / Kling AI team
Kling attracts attention after its public introduction in 2024 for motion and longer video generation. The master uses this node to show that generative-media development is not confined to European and North American laboratories. It also provides technical context for subsequent cross-platform disputes concerning audiovisual intellectual property.
Read original source · Kling video generation model ↗
Research master · p. 33 · #112 · Research claims await independent verification
Open in the interactive timeline ↗2024-06Dream Machine
+
Dream Machine
Luma AI
Dream Machine opens text- and image-to-video generation through a web product. The master compares the proliferation of video products in 2024 with the expansion of text-to-image products in 2022: users develop workflows around prompts, reference images and motion consistency as access becomes more widespread.
Read original source · Dream Machine ↗
Research master · p. 33 · #113 · Research claims await independent verification
Open in the interactive timeline ↗2024-06Gen-3 Alpha
+
Gen-3 Alpha
Runway
Gen-3 Alpha develops fidelity, motion and temporal consistency, with image and video controls added over time. The node marks the development of a more elaborate control layer for generative video, comparable to the specialized editing and conditioning tools already emerging around image generation.
Read original source · Gen-3 Alpha ↗
Research master · p. 33 · #114 · Research claims await independent verification
Open in the interactive timeline ↗2024-08FLUX.1 model family
+
FLUX.1 model family
Black Forest Labs (Robin Rombach; Andreas Blattmann; Patrick Esser et al.)
Black Forest Labs, founded with researchers previously involved in Stable Diffusion and CompVis, introduces FLUX.1. The family develops prompt adherence, text generation and a broader ecosystem for model use and adaptation. Community fine-tuning expands beyond SDXL toward another model lineage, subject to the different licenses of individual variants.
Read original source · FLUX.1 model family ↗
Research master · p. 36 · #124 · Research claims await independent verification
Open in the interactive timeline ↗2024-08Ideogram 2.0
+
Ideogram 2.0
Ideogram
Ideogram 2.0 develops typography, style control and multiple generation modes. The master situates this within a shift after 2024 from competition centered chiefly on photorealism toward readable text, brand character and practical design use. Specialized outputs become a stronger basis for differentiating image systems.
Read original source · Ideogram 2.0 ↗
Research master · p. 36 · #125 · Research claims await independent verification
Open in the interactive timeline ↗2024-10Stable Diffusion 3.5
+
Stable Diffusion 3.5
Stability AI
Stable Diffusion 3.5 offers Large, Large Turbo and Medium variants, continuing the MMDiT and flow-based approach while adjusting quality and hardware requirements. It belongs to a more plural ecosystem of accessible generative models, rather than reproducing the central community position once held by Stable Diffusion 1.5.
Read original source · Stable Diffusion 3.5 ↗
Research master · p. 36 · #126 · Research claims await independent verification
Open in the interactive timeline ↗2024-11FLUX.1 Tools
+
FLUX.1 Tools
Black Forest Labs
Black Forest Labs introduces Fill, Depth, Canny and Redux, incorporating editing, depth and edge guidance, and reference remixing into the model family. The node shows modular control moving from community extensions into model providers' own offerings. It records the historical release; current support for individual tools may differ.
Read original source · FLUX.1 Tools ↗
Research master · p. 36 · #127 · Research claims await independent verification
URL cited in the master (may have moved or expired) ↗
Open in the interactive timeline ↗2024-12HunyuanVideo
+
HunyuanVideo
Tencent Hunyuan team
HunyuanVideo releases model weights and code, extending competition over open access into large video-generation systems. It supports community efforts to arrange video models locally within tools such as ComfyUI. Availability of weights and code should be distinguished from any particular license's conditions of use.
Read original source · HunyuanVideo ↗
Research master · p. 33 · #115 · Research claims await independent verification
Open in the interactive timeline ↗2025-03Gen-4
+
Gen-4
Runway
Gen-4 emphasizes maintaining characters, environments and visual styles across generated shots. The master interprets this as a shift from isolated impressive clips toward production systems capable of supporting continuity, editing and sequential storytelling. These production ambitions also change what creators need to control and evaluate.
Read original source · Gen-4 ↗
Research master · p. 34 · #116 · Research claims await independent verification
Open in the interactive timeline ↗2025-03Gemini native image generation preview
+
Gemini native image generation preview
Google DeepMind / Gemini team
Gemini 2.0 Flash demonstrates native image capabilities through text-and-image dialogue, continuing edits and multimodal output. Alongside GPT-4o, it indicates image generation moving into general multimodal models, extending beyond separately operated diffusion products and changing how visual tasks are discussed and revised.
Read original source · Gemini native image generation preview ↗
Research master · p. 37 · #129 · Research claims await independent verification
Open in the interactive timeline ↗2025-03Ideogram 3.0
+
Ideogram 3.0
Ideogram
Ideogram 3.0 continues to develop text rendering, photorealism and style references. The master reads it as part of increasing specialization among image products in 2025, with different systems distinguishing themselves through design tasks, typography and the production of brand assets.
Read original source · Ideogram 3.0 ↗
Research master · p. 37 · #130 · Research claims await independent verification
Open in the interactive timeline ↗2025-03-25GPT-4o native image generation
+
GPT-4o native image generation
OpenAI image generation and multimodal teams
GPT-4o image generation works within conversational context to interpret uploaded images, render text, revise outputs over multiple turns and draw on model knowledge. The master calls this an Interface Turn: creators collaborate with a multimodal system that combines seeing, discussion, editing and generation, rather than operating an isolated image generator.
Read original source · GPT-4o native image generation ↗
Research master · p. 37 · #131 · Research claims await independent verification
Open in the interactive timeline ↗2025-04-04Midjourney V7
+
Midjourney V7
Midjourney team
V7 develops text and image prompting and detail consistency, while making personalization a default capability at launch. Conversational iteration in Draft Mode illustrates a shift even within dedicated image platforms from isolated prompts toward an ongoing workflow of discussion and revision.
Read original source · Midjourney V7 ↗
Research master · p. 38 · #132 · Research claims await independent verification
Open in the interactive timeline ↗2025-04-23gpt-image-1 API
+
gpt-image-1 API
OpenAI image generation team
OpenAI makes its image-generation capabilities available to developers through gpt-image-1. The historical point is a further movement of the interface: native multimodal image generation can be embedded in commerce, design, education, games and organizational workflows, instead of being available only inside a particular chat product.
Read original source · gpt-image-1 API ↗
Research master · p. 38 · #133 · Research claims await independent verification
Open in the interactive timeline ↗2025-04-24Firefly Image Model 4
+
Firefly Image Model 4
Adobe Firefly / Creative Cloud teams
Firefly Image Model 4 appears alongside partner models within a unified platform. Firefly increasingly functions as a layer for creative production rather than only as one generative model. The master compares the task-specific invocation of generative capabilities in professional software with the use of fonts, plug-ins or renderers.
Read original source · Firefly Image Model 4 ↗
Research master · p. 38 · #134 · Research claims await independent verification
Open in the interactive timeline ↗2025-05Veo 3
+
Veo 3
Google DeepMind
Veo 3 combines video and audio generation, reducing the need to assemble outputs from separate systems. The unit of generative media becomes a clip with sound rather than only an image or sequence of frames. This brings the technology into closer relation with audiovisual labor and chains of rights.
Read original source · Veo 3 ↗
Research master · p. 34 · #117 · Research claims await independent verification
Open in the interactive timeline ↗2025-05-20Imagen 4
+
Imagen 4
Google DeepMind / Imagen team
Imagen 4 is announced alongside Veo 3 and Lyria 2. The master sees this as evidence that image, video and audio generation are being organized as a collection of media capabilities within one platform, rather than remaining separate model categories with isolated interfaces and workflows.
Read original source · Imagen 4 ↗
Research master · p. 38 · #135 · Research claims await independent verification
Open in the interactive timeline ↗2025-05-29FLUX.1 Kontext
+
FLUX.1 Kontext
Black Forest Labs research team
FLUX.1 Kontext extends text-to-image generation into a process that accepts both images and text, maintaining relationships among characters and scenes through successive edits. The master identifies a growing emphasis on reference-led and editing-centered models, beyond generation that begins only from a written prompt.
Read original source · FLUX.1 Kontext ↗
Research master · p. 39 · #136 · Research claims await independent verification
Open in the interactive timeline ↗2025-08Gemini 2.5 Flash Image
+
Gemini 2.5 Flash Image
Google DeepMind / Gemini team
Gemini 2.5 Flash Image develops the preservation of people and objects, combination of multiple images and conversational editing. It represents the growing importance of generation grounded in references during 2025: creative work increasingly involves continuing to work with existing visual subjects, rather than always producing an image from scratch.
Read original source · Gemini 2.5 Flash Image ↗
Research master · p. 39 · #137 · Research claims await independent verification
Open in the interactive timeline ↗2026-04-21ChatGPT Images 2.0; GPT-Image-2 API
+
ChatGPT Images 2.0; GPT-Image-2 API
OpenAI image generation team
The master places ChatGPT Images 2.0 and GPT-Image-2 within a move toward production workflows, emphasizing instruction following, layouts, text and higher-resolution output. It also records availability through developer and coding interfaces. The official announcement confirms the release date; the wider interpretation and individual capability claims require separate assessment.
Read original source · ChatGPT Images 2.0; GPT-Image-2 API ↗
Research master · p. 39 · #138 · Release date checked; other claims pending
Open in the interactive timeline ↗2026-04-21GPT-Image-2 / ChatGPT Images 2.0
+
GPT-Image-2 / ChatGPT Images 2.0
OpenAI image generation team
The master records GPT-Image-2's introduction through the API and Codex, with an emphasis on complex visual tasks, layouts, text, instruction following and output up to 2K. Its interpretation is a shift from producing an attractive image toward generating assets for products and design systems. Current model documentation is linked separately from the historical account.
Read original source · GPT-Image-2 / ChatGPT Images 2.0 ↗
Research master · p. 39 · #139 · Research claims await independent verification
Open in the interactive timeline ↗2026-09-08ChatGPT Images 2.5 / GPT-Image-2.5 Flare & Sunburst
+
ChatGPT Images 2.5 / GPT-Image-2.5 Flare & Sunburst
OpenAI image generation team
The announcement presents improvements in detail, subject preservation, multi-turn editing and speed, alongside Sketch, Templates and comments placed directly on images. It introduces Flare and Sunburst API variants. The master reads this as a convergence of generation, reference, annotation, editing, templates and collaboration within the image-making interface of 2026.
Read original source · ChatGPT Images 2.5 / GPT-Image-2.5 Flare & Sunburst ↗
Research master · p. 40 · #140 · Release date checked; other claims pending
Open in the interactive timeline ↗Further reading & citation
This page is a selection and reading path, not an exhaustive history. Cite the relevant record and original source for specific claims. To cite this guide, give its title, URL and access date.
Explore the full event chronology ↗