ANTI-AI ARCHIVE / Research guides ANTI-AI ARCHIVE / RESEARCH GUIDES
Catalogue Selected entry Research guide
ANTI-AI ARCHIVE GUIDE
ARCHIVE TITLE CARD · NOT AN ORIGINAL COVER
AI art history is more than a sequence of image-model releases. This reading path brings theory, artists’ practices, exhibition institutions and technical tools together to ask who sets the rules, who participates in making a work, and how that work is exhibited and interpreted.
Read the research guide Click an entry to preview; double-click to read. Read guide & citation ↗
TOPIC READING ROOM
AI art history 225 / 225 Explore 222 nodes connecting AI art’s intellectual contexts, artistic practices, exhibitions and technologies, with bilingual research accounts and original sources.
— AI art history: historical threads, 1943–2026 AI art history is more than a sequence of image-model releases. This reading path brings theory, artists’ practices, exhibition institutions and technical tools together to ask who sets the rules, who participates in making a work, and how that work is exhibited and interpreted. Research guide · ANTI-AI ARCHIVE Read AI art history: historical threads, 1943–2026 — Why begin in 1943? 1943 is an intellectual-context marker selected by the research master, not a claim that AI art began in that year. Early neural computation, cybernetics, computer art and later machine learning are not interchangeable. Their differences matter; this genealogy does not treat them as a single, inevitable progression toward present-day generative AI. Research guide · ANTI-AI ARCHIVE Read Why begin in 1943? — How can you use these 222 nodes? Start with a period, then compare original titles, people and institutions, research accounts and source links. A theoretical paper can document a proposed method; an exhibition record can document display and circulation. Neither alone establishes a technology’s influence on all artistic practice. Each entry retains date precision, source limitations and verification status. The timeline and card catalogue scroll together. Use Read or double-click a card to open its dedicated page with the full account, sources and citation. Research guide · ANTI-AI ARCHIVE Read How can you use these 222 nodes? 1943 A Logical Calculus of the Ideas Immanent in Nervous Activity McCulloch and Pitts use simplified neurons and logical calculus to examine how neural activity could perform computation. This is not a paper on image generation. Its importance to this genealogy is the proposition that perception and cognition can be represented as computable networks: a formal prehistory that should remain visible when discussing neural image generation. Historical node · McCulloch & Pitts Read A Logical Calculus of the Ideas Immanent in Nervous Activity 1948 A Mathematical Theory of Communication Shannon establishes a mathematical theory of communication. The genealogy does not position him as the direct origin of AI art; it identifies a broader condition in which images, writing and sound can be treated as encodable information. Datasets, compressed representations, latent spaces and generative models develop within this wider informational framework. Historical node · Claude Shannon Read A Mathematical Theory of Communication 1948 Cybernetics Wiener's Cybernetics brings communication, feedback and control into one theoretical framework. For the archive's later questions of control and sovereignty, cultural information is not merely stored or transmitted: it can enter feedback loops and become a resource for regulating machine behavior and organising power. This connection is the master's historical interpretation. Historical node · Norbert Wiener Read Cybernetics 1949 The Organization of Behavior Hebb's account of learning emphasises that connections between neurons can change through joint activity. Although far removed from image generation, it helps establish a consequential hypothesis: capabilities need not be written into a system entirely as explicit rules, but may be acquired through experience and changes in connection strengths. Historical node · Donald Hebb Read The Organization of Behavior 1950 Computing Machinery and Intelligence Rather than first defining the essence of thought, Turing's imitation game reframes machine intelligence as a question of behavior and interaction. The master connects this operational shift with later assessments of visual generation through what systems can produce, modify and organize, rather than through a prior resolution of whether machines possess creativity. Historical node · Alan Turing Read Computing Machinery and Intelligence 1952 Oscillons / Electronic Abstractions Laposky generates electronic waveforms with an oscilloscope and records them photographically as abstract images. Before digital computer art becomes established, this practice connects electrical signals, mathematical curves and visual composition. It forms a prehistory of electronic generative aesthetics rather than an instance of contemporary machine learning. Historical node · Ben F. Laposky Read Oscillons / Electronic Abstractions 1954 Aesthetica / Introduction to Information-Theoretic Aesthetics Bense publishes four volumes of Aesthetica in 1954–1960, followed by an expanded collected edition in 1965 and an introduction to information-theoretic aesthetics in 1969. These works connect aesthetics, information theory, statistical structure and programmed production. The genealogy marks a turn toward computational descriptions of form, order, complexity and aesthetic difference; it does not assert that art is simply information. Historical node · Max Bense Read Aesthetica / Introduction to Information-Theoretic Aesthetics 1956 Dartmouth Summer Research Project on Artificial Intelligence proposal/workshop The Dartmouth workshop does not directly concern art. It establishes an academic community around the proposition that intelligence can be simulated and studied by machines. Later computer vision, machine learning and generative models develop within this institutional lineage. The node therefore provides disciplinary context for art history rather than documenting an artistic event. Historical node · John McCarthy; Marvin Minsky; Nathaniel Rochester; Claude Shannon and participants Read Dartmouth Summer Research Project on Artificial Intelligence proposal/workshop 1956 Oscillograms; early computer / electronic art writing Franke makes electronic oscillographic images and writes extensively about the roles of mathematics, information and technology in art. Alongside the information-aesthetic milieu associated with Bense and Nake, his work helps establish networks of theory and practice in early computer art in Germany and Austria. Historical node · Herbert W. Franke Read Oscillograms; early computer / electronic art writing 1958 The Perceptron Rosenblatt's perceptron links pattern recognition to neural-network learning. In the history of visual AI, it marks an early step from systems whose rules are entirely specified by programmers toward systems that adjust their parameters using examples. Its relevance here concerns this change in how visual capabilities are constructed. Historical node · Frank Rosenblatt Read The Perceptron 1958 Information Theory and Esthetic Perception Moles's French study later appears in English as Information Theory and Esthetic Perception. It examines how art objects are perceived as informational messages and how redundancy and complexity contribute to aesthetic experience. Alongside Bense's work, it supplies an important theoretical background for information aesthetics and subsequent debates about computational form. Historical node · Abraham Moles Read Information Theory and Esthetic Perception 1959 Some Studies in Machine Learning Using the Game of Checkers Samuel's checkers program demonstrates how a machine can improve its performance through learning without a programmer specifying every strategy step by step. The master situates this shift from explicit rules toward learning from experience within the longer technical history of data-driven visual systems and contemporary generative models. Historical node · Arthur Samuel Read Some Studies in Machine Learning Using the Game of Checkers 1960 Man-Computer Symbiosis Licklider proposes human–computer symbiosis: people and computers undertake different tasks and form a new cognitive arrangement through real-time interaction. The master relates this history of interaction to contemporary conversational image generation and visual agents. The connection concerns interfaces for collaborative work, rather than an identity between the earlier proposal and present models. Historical node · J. C. R. Licklider Read Man-Computer Symbiosis 1961 New Tendencies exhibitions and Computers and Visual Research New Tendencies in Zagreb moves from concrete and optical art toward Computers and Visual Research, placing programs, information, scientific methods and viewers' perception at the center of artistic discussion. The series provides an early international institutional platform for computer art, extending the genealogy beyond individual inventions or isolated laboratories. Historical node · Almir Mavignier · Matko Meštrović · international participants Read New Tendencies exhibitions and Computers and Visual Research 1963 Sketchpad Sketchpad brings graphical objects, constraints and human–computer interaction into one system, helping establish interactive computer graphics. It reminds us that a history of AI art concerns not only how models generate images, but also how interfaces turn computational capacities into usable visual practices. Historical node · Ivan Sutherland Read Sketchpad 1964 Human or Machine: A Subjective Comparison of Piet Mondrian’s Composition with Lines and a Computer-Generated Picture Noll generates computer images resembling a Mondrian composition and uses viewer experiments to compare aesthetic preferences and judgments of authorship. This anticipates later debates about machine imitation of style and the recognition of AI images. The technical mechanism, however, remains programmed generation rather than learning from training images. Historical node · A. Michael Noll Read Human or Machine: A Subjective Comparison of Piet Mondrian’s Composition with Lines and a Computer-Generated Picture 1965 Early computer-generated art exhibitions and Noll experiments In 1965, Nees exhibits computer graphics in Stuttgart; Noll and Béla Julesz show computer-generated pictures at New York's Howard Wise Gallery; Nake is also among the earliest public exhibitors of algorithmic art. The computer enters art institutions as a means of producing form, extending its role beyond engineering. Historical node · Georg Nees; Frieder Nake; A. Michael Noll Read Early computer-generated art exhibitions and Noll experiments 1966 9 Evenings: Theatre & Engineering 9 Evenings organizes artists and Bell Labs engineers into large-scale live experiments and contributes directly to the formation of E.A.T. The event expands the roles of computation and control systems in art, providing context for later interactive art, machine behavior and real-time generation. Historical node · Billy Klüver; Robert Rauschenberg; John Cage; Lucinda Childs; engineers and artists Read 9 Evenings: Theatre & Engineering 1966 Experiments in Art and Technology (organization) Billy Klüver, Fred Waldhauer, Robert Rauschenberg and Robert Whitman help establish Experiments in Art and Technology, institutionalizing collaboration between artists and engineers. E.A.T. is not an AI art organisation, but its collaborative production model becomes important to subsequent digital, media and experimental technological art. Historical node · Billy Klüver; Fred Waldhauer; Robert Rauschenberg; Robert Whitman Read Experiments in Art and Technology (organization) 1968 Computer Arts Society Founded in Britain, the Computer Arts Society connects artists, designers, engineers and researchers. Its significance is institutional: computer art begins to acquire sustained structures for exhibition, discussion, publication and professional exchange, moving beyond scattered experiments into a continuing field of practice. Historical node · Alan Sutcliffe; George Mallen; John Lansdown Read Computer Arts Society 1968 Cybernetic Serendipity exhibition Curated by Jasia Reichardt, Cybernetic Serendipity brings together computer music, computer graphics, machines and environments, and computer poetry. It moves the question of computers and art from laboratories into a broader public cultural setting, while placing cybernetics and artistic practice alongside one another. Historical node · Jasia Reichardt (curator); participating artists/researchers Read Cybernetic Serendipity exhibition 1968 Algorithmic drawings / “machine imaginaire” practice From a manually conceived machine imaginaire to actual computer programs, Molnár uses rules, permutations and small random deviations in painting and drawing. Her practice demonstrates that generative art was not suddenly invented by deep learning: artists had long treated procedures, constraints and variation as mechanisms through which works take form. Historical node · Vera Molnár Read Algorithmic drawings / “machine imaginaire” practice 1969 AARON Cohen begins conceiving AARON in the late 1960s and names and develops it during the 1970s. Instead of learning from a massive image dataset, AARON encodes his knowledge of drawing, representation and composition as rules. It offers an important comparison: machine art need not depend on scraping or style appropriation, so contemporary anti-AI politics cannot simply be equated with opposition to all machine-made art. Historical node · Harold Cohen Read AARON 1969 Computer-generated Drawings / A Programmed Aesthetics Mohr systematically uses computers and plotters after 1969 and holds a significant solo exhibition at the Musée d'Art Moderne de la Ville de Paris in 1971. His work demonstrates how a program can become an enduring system of authorship, comparable in its continuity to an artist's sustained method of painting. Historical node · Manfred Mohr Read Computer-generated Drawings / A Programmed Aesthetics 1970 Software exhibition, Jewish Museum, New York Jack Burnham's Software exhibition brings information processing, conceptual systems and interactive processes into an art institution. It connects systems aesthetics, conceptual art and computational culture, providing an important context for understanding artworks through rules, protocols, programs and informational processes rather than exclusively as discrete objects. Historical node · Jack Burnham (curator); artists and technologists Read Software exhibition, Jewish Museum, New York 1971 Art and Technology program/exhibition LACMA's Art and Technology program gives artists access to corporate laboratories and industrial resources. It is not an AI art project. Its place in this genealogy is institutional: it offers an early model for later production relationships among artists, large technology companies and research organisations. Historical node · Maurice Tuchman (curator); participating artists and corporations Read Art and Technology program/exhibition 1975 Fractal Objects (1975) / The Fractal Geometry of Nature (1982) Mandelbrot's fractal research is not AI, but has substantial relevance to generative art: simple rules can produce complex, self-similar visual worlds. The spread of fractal imagery during the 1980s and 1990s also helps connect mathematical computation and visual invention in public culture. Historical node · Benoît Mandelbrot Read Fractal Objects (1975) / The Fractal Geometry of Nature (1982) 1979 Ars Electronica Festival Beginning in 1979, Ars Electronica places art, technology and social questions on a shared public platform. Its continuing presentation of artificial life, robotics, data art and AI art provides long-term cultural infrastructure for practices that otherwise might remain isolated technological experiments. Historical node · Hannes Leopoldseder; Hubert Bognermayr; Herbert W. Franke; others Read Ars Electronica Festival 1986 Biomorph evolutionary program Dawkins's Biomorph program uses recursive rules and human selection to demonstrate cumulative evolution. It is not machine-learning art, but belongs alongside the evolutionary practices of Sims and Latham in a history of interactive evolutionary aesthetics, where creators select descendants within a generative space instead of drawing each result directly. Historical node · Richard Dawkins Read Biomorph evolutionary program 1986 Learning representations by back-propagating errors Backpropagation allows hidden layers to learn internal representations useful for a task. It is not itself a technique for visual generation, but becomes one of the mechanisms underlying later deep-learning vision systems. Features need no longer be entirely designed by hand; they can form through training. Historical node · Rumelhart, Hinton & Williams Read Learning representations by back-propagating errors 1987 Boids Boids simulates collective motion using local rules such as separation, alignment and cohesion. Its influence extends across animation, artificial life and generative art. The system shows how visual complexity can emerge from distributed rules rather than from a central designer specifying the behavior of the whole. Historical node · Craig Reynolds Read Boids 1988 FormGrow / Mutator evolutionary systems Latham and Todd develop evolutionary systems for generating three-dimensional forms in settings including IBM UK. Their approach shifts artistic work from direct shaping toward defining generative rules and selecting mutations. The master treats this as a historical analogy for contemporary latent-space exploration and workflows built around variation. Historical node · William Latham; Stephen Todd Read FormGrow / Mutator evolutionary systems 1991 Artificial Evolution for Computer Graphics Sims uses genetic algorithms and interactive selection to generate complex images, textures and animation. The machine produces candidates and users guide the next round through selection. This places aesthetic judgment inside the generative loop, rather than reserving it for evaluation after an artwork has been completed. Historical node · Karl Sims Read Artificial Evolution for Computer Graphics 1994 Evolving Virtual Creatures Sims extends generation beyond still images by evolving the body structures and control networks of virtual creatures together. Generative practice encompasses dynamic behavior and artificial life, providing context for later work in games, simulation art and embodied generative systems. Historical node · Karl Sims Read Evolving Virtual Creatures 1995 Pyramid-Based Texture Analysis/Synthesis; Parametric Texture Model These texture-synthesis methods reconstruct appearance through multiscale filter responses and statistical constraints. Gatys's later neural style transfer uses statistics of deep-network features to represent style. The earlier tradition of statistical texture analysis is therefore an important technical antecedent, without implying that aesthetic style is exhausted by such statistics. Historical node · David Heeger; James Bergen; Javier Portilla; Eero Simoncelli Read Pyramid-Based Texture Analysis/Synthesis; Parametric Texture Model 1998 Gradient-Based Learning Applied to Document Recognition LeNet-5 and related work integrate convolution, shared weights and gradient-based training into a stable visual architecture. Later developments including AlexNet, convolutional encoders in generative models and visual feature extraction inherit elements of this tradition. The node documents technical conditions for subsequent visual learning. Historical node · LeCun et al. Read Gradient-Based Learning Applied to Document Recognition 2001 Image Quilting for Texture Synthesis and Transfer Image Quilting selects and joins patches from sample textures to synthesize larger textures, and also demonstrates texture transfer. It is not deep learning. It documents an important earlier route for reconstructing visual appearance from existing image material, preceding the widespread use of neural generative models. Historical node · Alexei A. Efros; William T. Freeman Read Image Quilting for Texture Synthesis and Transfer 2001 Processing Processing makes drawing, interaction and program structure accessible within a creative-coding environment. Although not AI, it becomes important infrastructure for generative art in the 2000s. The master relates its culture of programmable visual practice to later use of notebooks, node workflows and prompt scripting. Historical node · Casey Reas; Ben Fry Read Processing 2006 A Fast Learning Algorithm for Deep Belief Nets This work is an important node in the revival of deep learning during the 2000s. It helps renew research into learning multilayer representations and creates conditions for later large-scale visual deep learning. Its historical role concerns representation learning rather than the direct production of artworks. Historical node · Hinton, Osindero & Teh Read A Fast Learning Algorithm for Deep Belief Nets 2006 The Painting Fool The Painting Fool is a long-running computational-creativity project that emphasises a system's creative decisions and accounts of its own process. It connects discussions of symbolic and computational creativity with the cultural questions later raised by data-driven generative models. Historical node · Simon Colton Read The Painting Fool 2009 ImageNet: A Large-Scale Hierarchical Image Database ImageNet connects WordNet's semantic hierarchy with millions of images and uses crowdsourcing for annotation. The master interprets this as a Dataset Turn: cultural and everyday images become organised as scalable resources for machine learning and comparison, alongside their existence as things people view. This turn is an editorial framework, not a neutral period label. Historical node · Jia Deng; Wei Dong; Richard Socher; Li-Jia Li; Kai Li; Li Fei-Fei Read ImageNet: A Large-Scale Hierarchical Image Database 2010 ILSVRC ILSVRC standardizes subsets of ImageNet into an annual visual-recognition competition. Data become not only training material, but a framework for measuring and comparing research performance. AlexNet's 2012 results are recognized and amplified within this institutional structure of datasets, benchmarks and competitive evaluation. Historical node · ImageNet team; international computer vision community Read ILSVRC 2012 AlexNet AlexNet substantially outperforms established methods on ImageNet classification and helps make deep convolutional networks central to computer vision. The master reads this as a convergence of dataset scale, GPU computation and learned representations, extending the Dataset Turn into the production of new visual capabilities. Historical node · Krizhevsky, Sutskever & Hinton Read AlexNet 2013 Auto-Encoding Variational Bayes The variational autoencoder establishes a differentiable framework for generative latent-variable modelling. For this genealogy, a significant legacy is the movement of latent space from statistical terminology into visual culture: images can be compressed into representations, reconstructed, and newly sampled within a learned generative space. Historical node · Kingma & Welling Read Auto-Encoding Variational Bayes 2014 Generative Adversarial Nets GANs organize generation through learning data distributions. In contrast with a rule-based system such as AARON, the logic producing an image is learned from training data rather than mainly being specified by an artist. The master identifies a change in how machine images are understood: outputs are treated as samples from a learned distribution. Historical node · Goodfellow et al. Read Generative Adversarial Nets 2014 Conditional Generative Adversarial Nets Conditional GANs supply additional information to both generator and discriminator, offering a general framework for controlled generation. Class-conditioned synthesis, text-to-image GANs and forms of structurally conditioned generation build on this approach. The node concerns the development of conditions that guide what a model generates. Historical node · Mehdi Mirza; Simon Osindero Read Conditional Generative Adversarial Nets 2015 LAPGAN LAPGAN generates image detail at multiple scales using a Laplacian pyramid. It is an important attempt at higher-resolution synthesis before Progressive GAN. The work illustrates a continuing problem in image generation: how to extend quality through the coordinated construction of images across layers or scales. Historical node · Emily Denton; Soumith Chintala; Arthur Szlam; Rob Fergus Read LAPGAN 2015 A Neural Algorithm of Artistic Style Neural style transfer represents content through deep convolutional features and describes style through feature statistics, then synthesizes a new image. It does not establish that artistic style is fundamentally identical to those statistics. Its cultural importance lies in making the extraction, computation and transfer of style appear tangible within visual practice. Historical node · Gatys, Ecker & Bethge Read A Neural Algorithm of Artistic Style 2015 DeepDream DeepDream amplifies internal network activations to make recognized patterns visible. It is not contemporary text-to-image generation, but gives a broad public an early encounter with a recognizable visual aesthetic arising from neural representations. Artists, media and online communities quickly incorporate that aesthetic into their practices and discussions. Historical node · Alexander Mordvintsev; Christopher Olah; Mike Tyka / Google Read DeepDream 2015 DRAW DRAW combines a variational autoencoder, recurrent networks and differentiable attention, generating images through successive operations that read from and write to a canvas. Although it does not become the dominant architecture of later products, it demonstrates an early combination of attention and iterative image generation. Historical node · Karol Gregor; Ivo Danihelka; Alex Graves; Danilo Rezende; Daan Wierstra Read DRAW 2015 DCGAN DCGAN establishes architectural practices adopted by many later visual GANs and demonstrates that directions in latent space can correspond to semantic changes. It brings manipulation of latent representations closer to a practical creative workflow, in which users explore meaningful variations instead of selecting only among unrelated generated images. Historical node · Radford, Metz & Chintala Read DCGAN 2016 pix2pix pix2pix learns from paired images to generate a target image from inputs such as label maps, edge maps and sketches. It extends generation beyond sampling from noise toward transformation conditioned on visual structure. This provides a clear technical lineage for later editing, redrawing and structural control. Historical node · Isola et al. Read pix2pix 2016 Prisma mobile application Applications such as Prisma turn neural style transfer into a visual interface accessible to ordinary users. The historical point here is not a new algorithm, but a change in use: style becomes something that can be invoked as an everyday visual effect through a consumer application. Historical node · Prisma Labs Read Prisma mobile application 2016 StackGAN: Text to Photo-realistic Image Synthesis with Stacked GANs StackGAN uses a two-stage process that first turns a description into low-resolution structure and then refines it into a higher-resolution image. It is a significant branch of text-to-image research before DALL·E and Imagen, demonstrating that prompt-driven image generation has a technical history preceding diffusion models. Historical node · Han Zhang; Tao Xu; Hongsheng Li; Shaoting Zhang; Xiaogang Wang; Xiaolei Huang; Dimitris Metaxas Read StackGAN: Text to Photo-realistic Image Synthesis with Stacked GANs 2017 Progressive Growing of GANs Progressive GAN increases resolution layer by layer to improve training stability and image detail. Its high-resolution synthetic portraits become culturally striking, and the work leads into StyleGAN. The node connects an architectural strategy for scaling generation with the changing public visibility of synthetic faces. Historical node · Karras et al. Read Progressive Growing of GANs 2017 Attention Is All You Need The Transformer is initially a sequence-modelling architecture, not a paper on image generation. Later systems including CLIP, DALL·E, DiT, MMDiT and native multimodal models nevertheless draw directly or indirectly on attention and Transformer approaches. Its inclusion documents a cross-domain technical condition rather than an artwork or image model in itself. Historical node · Vaswani et al. Read Attention Is All You Need 2017 CycleGAN CycleGAN uses cycle consistency to learn translation between unpaired image domains, extending tasks such as changes of season, transformations between painting and photographic domains, and changes in object appearance. It helps develop the idea that a visual domain or style can be learned statistically and transformed into another. Historical node · Zhu et al. Read CycleGAN 2017 Neural Discrete Representation Learning VQ-VAE learns a discrete codebook that compresses high-dimensional inputs into combinable representations. DALL·E, VQGAN and several autoregressive image models subsequently draw on the idea of visual tokens. The node concerns the representation of images in units that can participate in learned sequences and generative processes. Historical node · Aaron van den Oord; Oriol Vinyals; Koray Kavukcuoglu Read Neural Discrete Representation Learning 2017 AttnGAN: Fine-Grained Text to Image Generation with Attentional GANs AttnGAN aligns words with visual regions through attention and a deep attentional multimodal similarity model, improving details corresponding to complex descriptions. It provides an important pre-diffusion example of the problems now discussed as adherence to prompts and alignment between textual and visual information. Historical node · Tao Xu; Pengchuan Zhang; Qiuyuan Huang; Han Zhang; Zhe Gan; Xiaolei Huang; Xiaodong He Read AttnGAN: Fine-Grained Text to Image Generation with Attentional GANs 2018 BigGAN BigGAN scales GAN training for class-conditioned ImageNet generation and improves image fidelity. It reinforces an observation that recurs in the foundation-model period: model size, data and computational scale can themselves contribute to generative quality. The historical argument concerns this scaling relationship, not an automatic equation of scale with artistic value. Historical node · Brock, Donahue & Simonyan Read BigGAN 2018 StyleGAN StyleGAN introduces a mapping network and a style-based generator that offer more intuitive control over visual attributes at different scales. It broadens the cultural influence of navigable latent spaces, supporting face synthesis, latent editing and AI portrait practices. The model's technical use of style should remain distinct from the full art-historical concept. Historical node · Karras, Laine & Aila Read StyleGAN 2018 Ganbreeder / Artbreeder Joel Simon's Ganbreeder develops into Artbreeder, turning GAN latent spaces into interfaces for browsing, mixing, inheriting and sharing visual forms. Before prompt culture becomes dominant, this is a significant example of latent culture, where interaction with generative models is organised around visual variation and social circulation. Historical node · Joel Simon; Artbreeder community Read Ganbreeder / Artbreeder 2018 Portrait of Edmond de Belamy Christie's sale of Portrait of Edmond de Belamy for $432,500 makes GAN art an international media subject. It also brings disputes over the provenance of code, authorship and the attribution of technical labour into view. The artwork's value chain includes models, software, an artistic collective and the institutions of the art market. Historical node · Obvious (Hugo Caselles-Dupré; Pierre Fautrel; Gauthier Vernier); Robbie Barrat lineage controversy Read Portrait of Edmond de Belamy 2019 AI: More than Human The Barbican exhibition places historical automata, machine learning, robotics and contemporary AI art within a shared narrative space. It shows AI art becoming an established subject of institutional exhibition around 2019, providing context for the controversies that follow the wider spread of generative image products in 2022. Historical node · Barbican curatorial team; international artists/researchers Read AI: More than Human 2019 Memories of Passersby I Klingemann presents and sells an installation that continuously generates images. The artwork is therefore not simply one AI-produced picture, but a working generative system and its potentially continuing output. This raises questions about the status of the art object, collecting and authorship beyond those posed by a single generated image. Historical node · Mario Klingemann Read Memories of Passersby I 2019 Generating Diverse High-Fidelity Images with VQ-VAE-2 VQ-VAE-2 combines hierarchical discrete latent variables with autoregressive priors to produce high-quality images for its period. It supplies technical background for approaches such as DALL·E, in which visual representations are tokenized and modeled within larger generative systems. Historical node · Ali Razavi; Aaron van den Oord; Oriol Vinyals Read Generating Diverse High-Fidelity Images with VQ-VAE-2 2019 SPADE paper / GauGAN demo SPADE converts semantic layouts into high-quality images, while NVIDIA's GauGAN demonstration presents its potential as an interactive creative tool. It anticipates later interest in layout control, ControlNet and design-oriented generation, where a user specifies spatial organisation rather than relying on text alone. Historical node · Taesung Park; Ming-Yu Liu; Ting-Chun Wang; Jun-Yan Zhu / NVIDIA Read SPADE paper / GauGAN demo 2019 Analyzing and Improving the Image Quality of StyleGAN StyleGAN2 addresses visual artifacts in StyleGAN and improves the mapping from latent representations to images. It strengthens techniques for inversion, editing and attribute manipulation, helping develop a workflow in which an image can be generated, mapped back into a model's representation and edited within that space. Historical node · Tero Karras; Samuli Laine; Miika Aittala; Janne Hellsten; Jaakko Lehtinen; Timo Aila Read Analyzing and Improving the Image Quality of StyleGAN 2020 Denoising Diffusion Probabilistic Models DDPM demonstrates high-quality image synthesis with diffusion probabilistic models. In contrast with a GAN's direct generation, diffusion constructs an image through an iterative process of reversing noise. This approach subsequently becomes a major technical basis for text-to-image systems after 2022. Historical node · Ho, Jain & Abbeel Read Denoising Diffusion Probabilistic Models 2020 Taming Transformers / VQGAN VQGAN combines convolutional encoding of local visual structure with a Transformer's capacity for broader modelling, producing high-quality discrete visual representations. It directly influences early VQGAN+CLIP art practices and contributes technical experience to later generative models operating within compressed latent spaces. Historical node · Esser, Rombach & Ommer Read Taming Transformers / VQGAN 2020 Jukebox: A Generative Model for Music Jukebox uses hierarchical VQ-VAEs and autoregressive priors to generate music and approximate vocals. It is distinct from later text-driven music platforms. Its relevance is cross-media: learning generative capacities from large collections of cultural recordings is already becoming a paradigm that extends beyond images. Historical node · Prafulla Dhariwal et al.; OpenAI Read Jukebox: A Generative Model for Music 2020 Language Models are Few-Shot Learners GPT-3 is not an image model, but helps elevate the prompt from ordinary input text into an interface for specifying many different tasks. The master connects later systems such as DALL·E 3, conversational image generation and language-driven editing to this broader change: natural language becomes a control layer for models. Historical node · Tom B. Brown et al.; OpenAI Read Language Models are Few-Shot Learners 2021 Botto Botto links the generation of candidate images with community voting and auctions in a continuing creative loop. It combines generative models, collective taste, token-based governance and the art market. This makes it a case for studying machine authorship alongside the organisation of cultural production through platforms and communities. Historical node · Mario Klingemann; Botto community Read Botto 2021 VQGAN+CLIP community workflow Practitioners combine VQGAN's image synthesis with CLIP's text–image similarity to guide generation using language. This is a distributed practical arrangement rather than one canonical paper. Prompt recipes, artists' names, seeds, notebooks and iteration settings become part of a shared creative vocabulary within the open-source notebook ecosystem. Historical node · Community practitioners; Katherine Crowson and open-source notebook ecosystem Read VQGAN+CLIP community workflow 2021 CLIP announced Natural-language supervision creates a shared image-text representation that soon becomes a control signal for community art workflows. This dates the initial announcement; the later paper is a separate node. Historical node · OpenAI Read CLIP announced 2021 DALL·E announced Large-scale text-to-image generation becomes a clearly legible product/research category before the diffusion wave. The announcement and paper publication are recorded separately. Historical node · OpenAI Read DALL·E announced 2021 Zero-Shot Text-to-Image Generation DALL·E uses an autoregressive Transformer to model text and image tokens together, demonstrating broad compositional capabilities without task-specific examples. It makes the idea of producing a new image from a sentence publicly legible as a generative paradigm, connecting research in token modelling with a new interface to images. Historical node · Aditya Ramesh; Mikhail Pavlov; Gabriel Goh; Scott Gray; Chelsea Voss et al. / OpenAI Read Zero-Shot Text-to-Image Generation 2021 Learning Transferable Visual Models From Natural Language Supervision CLIP learns image–text alignment from internet image–text pairs, allowing language to indicate visual concepts in a zero-shot setting. It becomes a key component of early prompt-driven art practices including VQGAN+CLIP and Disco Diffusion, while bringing web-scale image–text data into the center of generative culture. Historical node · Alec Radford et al. / OpenAI Read Learning Transferable Visual Models From Natural Language Supervision 2021 Diffusion Models Beat GANs Architectural improvements and classifier guidance allow diffusion models to outperform prominent GAN approaches on ImageNet synthesis. The master identifies this as a technical turning point in the transition from GAN-centered image generation toward diffusion. The comparison concerns a particular research benchmark and period, rather than every possible artistic or technical application. Historical node · Dhariwal & Nichol Read Diffusion Models Beat GANs 2021 Disco Diffusion notebook/system From October 2021, Disco Diffusion develops a reusable Colab workflow combining diffusion, CLIP guidance, cutouts, image prompts and animation parameters. Its art-historical importance lies as much in practice as in foundational research: large numbers of users adopt prompt engineering and iterative parameter choices as forms of creative work. Historical node · Somnai; Gandamu; Disco Diffusion contributors/community Read Disco Diffusion notebook/system 2021 GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models GLIDE compares CLIP guidance with classifier-free guidance and demonstrates photorealistic text-to-image generation and inpainting. It is a precursor to DALL·E 2 and later diffusion-based products. The combination of generation and editing is especially relevant to the development of interfaces for modifying existing visual material. Historical node · Alex Nichol; Prafulla Dhariwal et al. / OpenAI Read GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models 2021 Latent Diffusion Models Latent diffusion operates in the compressed space of a pretrained autoencoder and accepts conditions such as text through cross-attention. It reduces computational costs and provides the direct technical foundation for Stable Diffusion, helping make an open text-to-image ecosystem practicable on consumer graphics hardware. Historical node · Rombach et al. Read Latent Diffusion Models 2022 LAION-5B: An open large-scale dataset for training next generation image-text models LAION-5B makes a large image–text pair index and filtering methods available, providing data infrastructure relevant to Stable Diffusion and related ecosystems. It also becomes central to disputes about training sources, artists' names and copyright. The master reads this as a shift in which the dataset itself becomes a political object; an index does not confer ownership of linked images. Historical node · Christoph Schuhmann et al. / LAION Read LAION-5B: An open large-scale dataset for training next generation image-text models 2022 Hierarchical Text-Conditional Image Generation with CLIP Latents / DALL·E 2 product DALL·E 2 combines CLIP representations with a diffusion-based image prior and decoder, improving realism and correspondence with text. Its research preview, beta and later access without a waitlist expand during 2022. The master emphasises this product transition as a step in the international popularization of text-to-image generation. Historical node · Aditya Ramesh; Prafulla Dhariwal; Alex Nichol; Casey Chu / OpenAI Read Hierarchical Text-Conditional Image Generation with CLIP Latents / DALL·E 2 product 2022 Civitai model-sharing platform Platforms such as Civitai turn fine-tuned models and configurations from the Stable Diffusion ecosystem into circulating cultural objects. Generative capability diversifies from a single foundation model into a large community market. This also complicates questions of provenance, style imitation, models of identifiable people and consent. Historical node · Civitai founders/team and model-sharing community Read Civitai model-sharing platform 2022 LoRA: Low-Rank Adaptation of Large Language Models + diffusion community adoption The original LoRA paper appears in 2021, reducing fine-tuning costs by freezing a main model and training low-rank update matrices. During 2022–2023, the Stable Diffusion community adapts this into small exchangeable modules for styles, characters, clothing, poses and likenesses. Style and identity can circulate as lightweight model files as well as names in prompts. Historical node · Edward Hu et al. (LoRA); Stable Diffusion open-source community Read LoRA: Low-Rank Adaptation of Large Language Models + diffusion community adoption 2022 Midjourney V1 default era An early Midjourney generation era establishes a distinctive painterly aesthetic and Discord- first social creation loop. The month identifies a default-model period, not an exact release day. Historical node · Midjourney Read Midjourney V1 default era 2022 Midjourney V2 Rapid model iteration begins on a monthly/quarterly cadence, foreshadowing the release tempo that will define GenAI visual products. Historical node · Midjourney Read Midjourney V2 2022 Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding Imagen uses a large text encoder and diffusion decoders. The research highlights gains in image–text alignment from scaling the language model. It develops a direction that later multimodal systems also pursue: improvements in understanding language can change the visual capabilities and controllability of image generation. Historical node · Chitwan Saharia et al. / Google Research Read Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding 2022 DALL·E 2 research preview expands OpenAI starts adding up to 1,000 waitlist users per week; model evaluation becomes mass user testing. The preview expansion is distinct from paid beta access and removal of the waitlist. Historical node · OpenAI Read DALL·E 2 research preview expands 2022 Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Parti explores an autoregressive approach alongside diffusion, generating image tokens as another kind of language. It shows that the expansion of text-to-image generation in 2022 did not follow a single technical route. Token-based modelling remains a parallel way to organize the relation between descriptions and visual outputs. Historical node · Jiahui Yu et al. / Google Research Read Scaling Autoregressive Models for Content-Rich Text-to-Image Generation 2022 Diffusers library Diffusers organizes schedulers, pipelines, model components and training scripts as reusable software modules, lowering barriers to deploying and extending diffusion models. It is an intermediate layer through which research models become ecosystems of applications, making software infrastructure an important part of the genealogy of generative practice. Historical node · Hugging Face Diffusers team Read Diffusers library 2022 Midjourney open beta / Discord platform Midjourney's historical importance extends beyond model quality. Its Discord interface lets participants see one another's prompts and results, supporting social learning, the circulation of styles and prompt imitation. Text-to-image generation becomes a public visual culture as well as an individual creative tool. Historical node · David Holz; Midjourney team/community Read Midjourney open beta / Discord platform 2022 DALL·E 2 beta with pricing Commercialized access shifts image generation from research preview to consumer creative service. The beta is distinct from the later removal of the waitlist. Historical node · OpenAI Read DALL·E 2 beta with pricing 2022 stable-diffusion-webui The WebUI brings text-to-image generation, image-to-image transformation, inpainting, samplers and extensions into one interface, later incorporating LoRA and ControlNet. Community software of this kind is essential to the practical transformation of downloadable model weights into tools that large numbers of creators can actually use. Historical node · AUTOMATIC1111 and open-source contributors Read stable-diffusion-webui 2022 An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion Textual Inversion learns a new text embedding without retraining an entire model. Visual concepts, objects and styles can be represented through a new composable word within the system. This is an effective method of personalisation, not a claim that the nature of artistic style can be reduced to a token. Historical node · Rinon Gal; Yuval Alaluf; Yuval Atzmon; Or Patashnik; Amit Bermano; Gal Chechik; Daniel Cohen-Or Read An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion 2022 Stable Diffusion public model release Stable Diffusion builds on latent diffusion and makes model weights publicly available. The master emphasises a change in access: users can download, fine-tune, fork, combine and run models locally, rather than only access a corporate API. Community systems including LoRA workflows, ControlNet, Civitai and ComfyUI develop within this broader ecosystem. Historical node · Robin Rombach; Patrick Esser; CompVis; Stability AI; Runway collaborators Read Stable Diffusion public model release 2022 DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation DreamBooth fine-tunes a pretrained text-to-image model using a small set of subject images, associating a unique identifier with that subject. It brings preservation of identity into the generative process and supplies technical context for later customization of people, characters and products, as well as disputes concerning likeness. Historical node · Nataniel Ruiz; Yuanzhen Li; Varun Jampani; Yael Pritch; Michael Rubinstein; Kfir Aberman Read DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation 2022 DALL·E 2 Outpainting Generative imaging expands beyond the frame, normalizing AI as an editing/extension tool rather than only an image generator. Historical node · OpenAI Read DALL·E 2 Outpainting 2022 Théâtre D’opéra Spatial (Space Opera Theater) Jason Allen enters and wins a competition with a work made using Midjourney and subsequent editing. The event becomes a cultural marker of what the master calls the 2022 Political Turn: discussion shifts from technical demonstrations toward creative labour, disclosure rules, authorship and fairness in artistic competition. Historical node · Jason M. Allen; Midjourney Read Théâtre D’opéra Spatial (Space Opera Theater) 2022 DALL·E 2 waitlist removed High-quality image generation becomes openly accessible to millions rather than invitation- only. This is a change in access, separate from the model announcement. Historical node · OpenAI Read DALL·E 2 waitlist removed 2022 Make-A-Video Text-to-video becomes a visible research frontier only months after mass text-to-image adoption. Historical node · Meta Read Make-A-Video 2022 Unsupervised Unsupervised uses data from MoMA's collection to train or drive a generative visual system presented on a large real-time screen. It makes relations among datasets, museum collections and generative models visible as artistic material. The institutional collection participates in both the system's input and the public interpretation of its output. Historical node · Refik Anadol Studio; MoMA Read Unsupervised 2022 InstructPix2Pix: Learning to Follow Image Editing Instructions InstructPix2Pix uses GPT-3 and Stable Diffusion to synthesize training examples, then trains a diffusion model conditioned on editing instructions. It shifts the interaction from making a picture in a single generation toward changing an existing image through language, connecting generative modelling with iterative visual editing. Historical node · Tim Brooks; Aleksander Holynski; Alexei A. Efros Read InstructPix2Pix: Learning to Follow Image Editing Instructions 2022 Midjourney V4 A new architecture substantially improves coherence and aesthetics, showing the speed of closed-model iteration. Historical node · Midjourney Read Midjourney V4 2022 DALL·E API public beta Image generation becomes programmable infrastructure for third-party products. It records developer access rather than a second model launch. Historical node · OpenAI Read DALL·E API public beta 2022 Stable Diffusion 2.0 New OpenCLIP text encoder, new training subset and higher resolutions show how model/data changes quickly alter style and prompting behavior. Historical node · Stability AI Read Stable Diffusion 2.0 2022 Text to Image / Magic Media precursor A mainstream design platform embeds text-to-image directly into everyday design workflows. Historical node · Canva Read Text to Image / Magic Media precursor 2022 Stable Diffusion 2.1 Only weeks after 2.0, Stability changes data filtering and prompt behavior, illustrating release cycles measured in weeks. Historical node · Stability AI Read Stable Diffusion 2.1 2023 ComfyUI ComfyUI makes a Stable Diffusion workflow explicit as a graph. Users arrange models and control modules within a visual system rather than only entering a prompt. The master interprets this as a shift from using a model toward designing a production pipeline, where organisation and reuse of processes become part of image-making. Historical node · ComfyAnonymous and open-source contributors Read ComfyUI 2023 AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning AnimateDiff introduces motion modelling as a pluggable component within personalized text-to-image diffusion systems. Existing character models, styles and community assets can participate in animated outputs. This extends generation from still images toward video while retaining parts of the earlier image-model ecosystem. Historical node · Yuwei Guo et al. Read AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning 2023 Adding Conditional Control to Text-to-Image Diffusion Models ControlNet freezes a pretrained diffusion backbone and learns spatial conditions through an additional network. It extends control from textual semantics toward geometry and composition. This is one technical condition for more specialized production workflows, in which artists or designers need to specify structure as well as subject matter. Historical node · Lvmin Zhang; Anyi Rao; Maneesh Agrawala Read Adding Conditional Control to Text-to-Image Diffusion Models 2023 T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models Alongside ControlNet, T2I-Adapter exemplifies a move toward fine-grained external control through modular attachments without fully retraining a foundation model. Image production becomes a combination of a base model and adapters. The node situates controllability in a growing ecosystem of interoperable components. Historical node · Chong Mou; Xintao Wang; Liangbin Xie; Yanze Wu; Jian Zhang; Zhongang Qi; Ying Shan; Xiaohu Qie Read T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models 2023 Gen-1 Video-to-video generation applies prompt/image style and composition to source footage, pushing GenAI toward filmmaking workflows. Historical node · Runway Read Gen-1 2023 Magic Design / expanded AI Visual Suite Generative AI moves from image creation into automated layouts, brand templates and end-to- end visual communication. Historical node · Canva Read Magic Design / expanded AI Visual Suite 2023 Midjourney V5 Photographic realism sharply improves and Midjourney becomes a reference point for high- aesthetic closed image generation. Historical node · Midjourney Read Midjourney V5 2023 Adobe Firefly / Generative Fill Firefly places generative AI inside established creative-software infrastructure. Its launch materials connect the product with questions about training sources, creator compensation, training preferences and content credentials. The master treats this as a meeting between model development and the governance of professional creative production; product commitments must be distinguished from independently demonstrated outcomes. Historical node · Adobe Firefly / Adobe Research and Creative Cloud teams Read Adobe Firefly / Generative Fill 2023 Bing Image Creator DALL·E-powered generation becomes embedded directly in search, chat and the Edge browser. Historical node · Microsoft Read Bing Image Creator 2023 Niji 5 Anime-specific model specialization becomes a mainstream product strategy rather than an informal fine-tune niche. Historical node · Midjourney / Spellbrush Read Niji 5 2023 SDXL beta SDXL marks the next open ecosystem backbone with better composition and photorealism. This beta access for API and DreamStudio users is separate from version 1.0. Historical node · Stability AI Read SDXL beta 2023 Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold DragGAN uses a GAN's feature space for point-based manipulation, demonstrating direct interaction with a generative image manifold. Although diffusion becomes increasingly prominent, this approach remains relevant to the development of editing interfaces in which users move elements visually instead of specifying every change through text. Historical node · Xingang Pan et al.; Max Planck Institute for Informatics / Saarland research team Read Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold 2023 Generative Fill in Photoshop beta Diffusion-based generation becomes a native layer inside the dominant professional raster- editing application. Historical node · Adobe Read Generative Fill in Photoshop beta 2023 Midjourney V5.1 Natural-language prompting and default aesthetics improve within weeks of V5, illustrating fast closed-model iteration. Historical node · Midjourney Read Midjourney V5.1 2023 Gen-2 multimodal video generation Runway Gen-2 supports short-video generation conditioned on text and images. It is an important example of video generation becoming a product in 2023. Creative concerns expand from a single composition toward motion, camera behavior, temporal consistency and the demands of audiovisual production. Historical node · Runway Read Gen-2 multimodal video generation 2023 Midjourney V5.2 Sharper detail, improved composition and stylize controls mature the V5 family. This node records refinements within the V5 model family. Historical node · Midjourney Read Midjourney V5.2 2023 CM3leon A token-based multimodal transformer shows competitive image generation and image-to-text within one foundation model. Historical node · RESEARCH MODEL Read CM3leon 2023 SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis SDXL develops composition, detail, style and resolution while retaining a model ecosystem that can be fine-tuned and combined with other tools. The master positions it within the Stable Diffusion community's movement from rapid experimentation toward more established production practices. Historical node · Robin Rombach; Stability AI research team and collaborators Read SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis 2023 IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models IP-Adapter uses decoupled cross-attention to add lightweight image-prompt capabilities to pretrained diffusion models. Reference-image control can coexist with text prompts and ControlNet, supporting workflows for portraits, fashion and identity that combine several types of guidance rather than relying on a description alone. Historical node · Hu Ye; Jun Zhang; Sibo Liu; Xiao Han; Wei Yang / Tencent AI Lab Read IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models 2023 Ideogram text-to-image model Ideogram attracts attention for rendering text within generated images. The node illustrates how text-to-image competition extends beyond visually appealing scenes toward typography, layouts, logos and commercial graphics. Design usability becomes a distinct area of model development and product differentiation. Historical node · Ideogram team Read Ideogram text-to-image model 2023 PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis PixArt-α focuses on diffusion Transformers, pretrained text encoders and efficient training. It is an open research node leading into the wider development of Transformer- and flow-based image systems in 2024. Its significance concerns both architectural choices and the resources needed to train visual models. Historical node · Junsong Chen et al. Read PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis 2023 DALL·E 3 / ChatGPT integration Integration with ChatGPT lets users discuss an idea and request changes while a language model organizes image prompts. The interface shifts from a prompt box toward conversation, reducing the need to master a particular model's prompting conventions. The genealogy follows this change in interaction as well as changes in image quality. Historical node · OpenAI image generation team; ChatGPT product team Read DALL·E 3 / ChatGPT integration 2023 DALL·E 3 announced Improved prompt following and text rendering are paired with LLM-mediated prompting rather than specialist prompt syntax. ChatGPT availability is recorded separately from this announcement. Historical node · OpenAI Read DALL·E 3 announced 2023 Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference Latent Consistency Models use consistency distillation in latent space to generate high-resolution images in a small number of steps, including two to four in the reported approach. Speed changes the possible interface: generation closer to real time can participate more readily in drawing, design, live presentation and interactive software. Historical node · Simian Luo et al. Read Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference 2023 Firefly Image 2 / Vector / Design Models Adobe expands from raster generation into vector graphics and editable design templates. Historical node · Adobe Read Firefly Image 2 / Vector / Design Models 2023 Pika 1.0 Pika packages text-to-video, image-to-video and video modification as accessible creative tools. Alongside Runway, it helps move generative video from research and specialist experimentation toward consumer products. Its historical role includes the interface and availability of the tool, as well as the underlying model. Historical node · Pika Labs Read Pika 1.0 2023 Stable Video Diffusion Stable Video Diffusion makes latent video diffusion available to researchers and developers, extending an open generative ecosystem from still images toward video. It also points toward increasingly shared infrastructure between image and video models, including software for arranging and adapting generation processes. Historical node · Stability AI / research collaborators Read Stable Video Diffusion 2023 Ideogram 0.2 A rapid follow-up improves prompting and typography only two months after launch. Historical node · Ideogram Read Ideogram 0.2 2023 Emu Video + Emu Edit Meta combines text-to-video with instruction-based image editing and begins tying both to a shared image foundation model. Historical node · Meta Read Emu Video + Emu Edit 2023 SDXL Turbo Adversarial Diffusion Distillation cuts generation to one step and turns real-time text-to-image into a practical product goal. Historical node · Stability AI Read SDXL Turbo 2023 Midjourney V6 Midjourney V6 develops prompt accuracy, coherence, image prompting and remix capabilities. It represents continuing iteration within a closed platform, where visual quality, aesthetic character and community interfaces evolve together. The node documents one product lineage alongside the open-weight and research-model histories elsewhere in the timeline. Historical node · Midjourney team Read Midjourney V6 2023 Imagen 2 rollout Imagen generation is integrated into Google Cloud/Vertex and consumer experiments, increasing enterprise availability. Historical node · Google Read Imagen 2 rollout 2024 Recraft V3 Design-oriented systems such as Recraft extend generation from visual imagination toward brand assets, typography, vectors and deliverable design material. This differentiation shows image generation becoming a set of tools for distinct creative industries, rather than remaining one undifferentiated technical category. Historical node · Recraft team Read Recraft V3 2024 InstantID: Zero-shot Identity-Preserving Generation in Seconds InstantID combines facial embeddings with structural conditions to support efficient identity preservation in diffusion models. It marks a move from training specifically on a person, as in DreamBooth-style workflows, toward generation guided by references. The technique is directly relevant to later questions of likeness and consent. Historical node · Qixun Wang et al. Read InstantID: Zero-shot Identity-Preserving Generation in Seconds 2024 Sora Sora shifts public expectations of generative models toward sustained scenes, camera motion and physical consistency. The master sees this as widening research on generative visual culture to encompass video models and film production. Expectations or claims about world simulation should be distinguished from independently demonstrated capabilities. Historical node · OpenAI Sora research and product teams Read Sora 2024 Stable Cascade Stable Cascade builds on the Würstchen approach, combining strong semantic compression with staged generation. Although it does not become the dominant community ecosystem, it documents exploration of efficiency, compressed representations and modular generation in the period following SDXL. Historical node · Stability AI; Würstchen research lineage Read Stable Cascade 2024 Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis Stable Diffusion 3 combines a Transformer-based architecture with flow matching and emphasises multi-subject prompts, text and image quality. It represents the further incorporation of Transformer approaches into mainstream image generation. The associated research paper provides a technical account distinct from the product announcement. Historical node · Stability AI research team Read Stable Diffusion 3 / Scaling Rectified Flow Transformers for High-Resolution Image Synthesis 2024 ImageFX Google exposes Imagen generation in a dedicated consumer experimentation interface. Historical node · Google Read ImageFX 2024 Ideogram 1.0 Typography and prompt fidelity mature into a serious design-oriented image product. Version 1.0 is distinct from the initial 0.1 launch in 2023. Historical node · Ideogram Read Ideogram 1.0 2024 Recraft V2 A model explicitly optimized for designers emphasises brand consistency, vectors and controllable graphic output. Historical node · Recraft Read Recraft V2 2024 Image Services API Generation, upscale, erase, outpaint and structure/style controls are packaged as production APIs rather than standalone models. Historical node · Stability AI Read Image Services API 2024 Stable Diffusion 3 API / Turbo SD3 reaches developers before open weights, showing increasing separation between research model, hosted API and local release. Historical node · Stability AI Read Stable Diffusion 3 API / Turbo 2024 Firefly Image 3 Photorealism, style controls and reference-based generation deepen GenAI integration across Photoshop and InDesign. Historical node · MODEL/PRODUCT Read Firefly Image 3 2024 Photoshop Generate Image + upgraded Generative Fill Professional image editing shifts from selective fill toward full blank-canvas generation plus reference controls. Historical node · Adobe Read Photoshop Generate Image + upgraded Generative Fill 2024 Veo video generation model Announced at Google I/O 2024, Veo is presented with capabilities including 1080p output, sequences longer than a minute and interpretation of cinematic language. The node records a shift in product ambitions from brief demonstrations toward questions of duration, camera work and control associated with audiovisual production. Historical node · Google DeepMind Read Veo video generation model 2024 Imagen 3 Imagen 3 is introduced at Google I/O and subsequently enters Gemini and Vertex AI. The master treats this as an example of general-purpose AI platforms incorporating image generation as a standard capability within their model stacks, rather than maintaining it solely as a separate experimental product. Historical node · Google DeepMind / Imagen team Read Imagen 3 2024 Kling video generation model Kling attracts attention after its public introduction in 2024 for motion and longer video generation. The master uses this node to show that generative-media development is not confined to European and North American laboratories. It also provides technical context for subsequent cross-platform disputes concerning audiovisual intellectual property. Historical node · Kuaishou / Kling AI team Read Kling video generation model 2024 Dream Machine Dream Machine opens text- and image-to-video generation through a web product. The master compares the proliferation of video products in 2024 with the expansion of text-to-image products in 2022: users develop workflows around prompts, reference images and motion consistency as access becomes more widespread. Historical node · Luma AI Read Dream Machine 2024 Gen-3 Alpha Gen-3 Alpha develops fidelity, motion and temporal consistency, with image and video controls added over time. The node marks the development of a more elaborate control layer for generative video, comparable to the specialized editing and conditioning tools already emerging around image generation. Historical node · Runway Read Gen-3 Alpha 2024 Niji 6 MODEL Anime-specialized generation improves Japanese text and fine illustrative detail. The date follows the cited official legacy documentation. Historical node · Midjourney / Spellbrush Read Niji 6 2024 Stable Diffusion 3 Medium open release SD3 weights arrive for self-hosting, bringing MMDiT/flow-matching into the open ecosystem. Historical node · Stability AI Read Stable Diffusion 3 Medium open release 2024 LivePortrait Portrait animation becomes highly efficient and controllable, accelerating talking-head and character-animation workflows. Historical node · Kuaishou / KwaiVGI Read LivePortrait 2024 Midjourney V6.1 Detail coherence improves and generation becomes about 25% faster while retaining the V6 interaction model. The speed comparison should be read in the context of the publisher’s test conditions. Historical node · Midjourney Read Midjourney V6.1 2024 FLUX.1 model family Black Forest Labs, founded with researchers previously involved in Stable Diffusion and CompVis, introduces FLUX.1. The family develops prompt adherence, text generation and a broader ecosystem for model use and adaptation. Community fine-tuning expands beyond SDXL toward another model lineage, subject to the different licenses of individual variants. Historical node · Black Forest Labs (Robin Rombach; Andreas Blattmann; Patrick Esser et al.) Read FLUX.1 model family 2024 Ideogram 2.0 Ideogram 2.0 develops typography, style control and multiple generation modes. The master situates this within a shift after 2024 from competition centered chiefly on photorealism toward readable text, brand character and practical design use. Specialized outputs become a stronger basis for differentiating image systems. Historical node · Ideogram Read Ideogram 2.0 2024 Firefly Video Model preview Adobe brings generative video into a rights-cleared professional workflow narrative tied to Premiere Pro. This describes the company’s positioning, not an independent resolution of all rights questions. Historical node · Adobe Read Firefly Video Model preview 2024 Stable Diffusion 3.5 Stable Diffusion 3.5 offers Large, Large Turbo and Medium variants, continuing the MMDiT and flow-based approach while adjusting quality and hardware requirements. It belongs to a more plural ecosystem of accessible generative models, rather than reproducing the central community position once held by Stable Diffusion 1.5. Historical node · Stability AI Read Stable Diffusion 3.5 2024 Act-One Facial performance transfer turns a driving performance into generated character animation, expanding control over acting. Historical node · Runway Read Act-One 2024 Dream Lab Canva replaces/extends generic image generation with Leonardo Phoenix-powered visual creation inside a mass-market design suite. Historical node · Canva / Leonardo.Ai Read Dream Lab 2024 Movie Gen Text-to-video, personalized video, video editing and synchronized audio are unified in one research suite. Historical node · Meta Read Movie Gen 2024 Generate Video beta / Generative Extend beta Generative video moves into Firefly web and Premiere timelines as an editing primitive, not only a standalone clip generator. Historical node · Adobe Read Generate Video beta / Generative Extend beta 2024 Canvas / Magic Fill / Extend Infinite-canvas editing consolidates generation, inpainting and outpainting into one visual workspace. Historical node · Ideogram Read Canvas / Magic Fill / Extend 2024 Stable Diffusion 3.5 Medium A smaller consumer-hardware variant broadens access one week after the 3.5 family launch. Historical node · OPEN MODEL Read Stable Diffusion 3.5 Medium 2024 FLUX.1 Tools Black Forest Labs introduces Fill, Depth, Canny and Redux, incorporating editing, depth and edge guidance, and reference remixing into the model family. The node shows modular control moving from community extensions into model providers' own offerings. It records the historical release; current support for individual tools may differ. Historical node · Black Forest Labs Read FLUX.1 Tools 2024 SD3.5 ControlNets Canny, Depth and Blur controls make the new Stable family usable in precise production workflows. Historical node · Stability AI Read SD3.5 ControlNets 2024 HunyuanVideo HunyuanVideo releases model weights and code, extending competition over open access into large video-generation systems. It supports community efforts to arrange video models locally within tools such as ComfyUI. Availability of weights and code should be distinguished from any particular license's conditions of use. Historical node · Tencent Hunyuan team Read HunyuanVideo 2024 Frames Runway expands beyond video with an image model emphasising aesthetic worlds and stylistic consistency. Historical node · Runway Read Frames 2024 Pika 2.0 The product shifts toward more controllable scene/subject workflows and viral effect-oriented creation. Historical node · Pika Read Pika 2.0 2024 Sora public product launch Sora transitions from research preview into a consumer creative product, making frontier video generation broadly testable. Historical node · OpenAI Read Sora public product launch 2024 Veo 2 Video quality, camera control and realism improve rapidly within seven months of the first Veo announcement. Historical node · Google DeepMind Read Veo 2 2024 Whisk Prompting becomes image-first: subject, scene and style references can be remixed without writing long text prompts. Historical node · Google Read Whisk 2025 Aleph video editing Generative video shifts toward editing existing footage through language rather than only synthesizing new clips. The master provides year precision only; an exact day has not been established. Historical node · Runway Read Aleph video editing 2025 Gen-4.5 Runway’s next video model focuses on sequenced instructions, camera choreography and fidelity, reflecting rapid intra-year model turnover. Year precision follows the master; the exact release date remains to be checked. Historical node · Runway Read Gen-4.5 2025 Pika 2.1 / 2.2 family Consumer video products iterate on longer clips, effects and greater control while competing through viral creation formats. The master groups these versions; their individual dates remain to be checked. Historical node · Pika Read Pika 2.1 / 2.2 family 2025 Ray2 A larger multimodal architecture with 10x compute improves motion coherence and production readiness. The compute and performance claims are attributed to the release account. Historical node · Luma AI Read Ray2 2025 Ideogram 2a A fast intermediate model variant shows model lines splitting by speed/cost as well as quality. Historical node · Ideogram Read Ideogram 2a 2025 Gen-4 Gen-4 emphasises maintaining characters, environments and visual styles across generated shots. The master interprets this as a shift from isolated impressive clips toward production systems capable of supporting continuity, editing and sequential storytelling. These production ambitions also change what creators need to control and evaluate. Historical node · Runway Read Gen-4 2025 Gemini native image generation preview Gemini 2.0 Flash demonstrates native image capabilities through text-and-image dialogue, continuing edits and multimodal output. Alongside GPT-4o, it indicates image generation moving into general multimodal models, extending beyond separately operated diffusion products and changing how visual tasks are discussed and revised. Historical node · Google DeepMind / Gemini team Read Gemini native image generation preview 2025 Ideogram 3.0 Ideogram 3.0 continues to develop text rendering, photorealism and style references. The master reads it as part of increasing specialization among image products in 2025, with different systems distinguishing themselves through design tasks, typography and the production of brand assets. Historical node · Ideogram Read Ideogram 3.0 2025 GPT-4o native image generation GPT-4o image generation works within conversational context to interpret uploaded images, render text, revise outputs over multiple turns and draw on model knowledge. The master calls this an Interface Turn: creators collaborate with a multimodal system that combines seeing, discussion, editing and generation, rather than operating an isolated image generator. Historical node · OpenAI image generation and multimodal teams Read GPT-4o native image generation 2025 Kling 2.0 Kling reaches a major second-generation release after roughly ten months and more than twenty iterations. Historical node · Kuaishou Read Kling 2.0 2025 Midjourney V7 V7 develops text and image prompting and detail consistency, while making personalisation a default capability at launch. Conversational iteration in Draft Mode illustrates a shift even within dedicated image platforms from isolated prompts toward an ongoing workflow of discussion and revision. Historical node · Midjourney team Read Midjourney V7 2025 gpt-image-1 API OpenAI makes its image-generation capabilities available to developers through gpt-image-1. The historical point is a further movement of the interface: native multimodal image generation can be embedded in commerce, design, education, games and organisational workflows, instead of being available only inside a particular chat product. Historical node · OpenAI image generation team Read gpt-image-1 API 2025 Firefly Image Model 4 Firefly Image Model 4 appears alongside partner models within a unified platform. Firefly increasingly functions as a layer for creative production rather than only as one generative model. The master compares the task-specific invocation of generative capabilities in professional software with the use of fonts, plug-ins or renderers. Historical node · Adobe Firefly / Creative Cloud teams Read Firefly Image Model 4 2025 Firefly Video Model GA + Firefly Boards Adobe turns Firefly into a multi-model creative studio spanning image, video, audio, vector and collaborative ideation. Historical node · Adobe Read Firefly Video Model GA + Firefly Boards 2025 Veo 3 Veo 3 combines video and audio generation, reducing the need to assemble outputs from separate systems. The unit of generative media becomes a clip with sound rather than only an image or sequence of frames. This brings the technology into closer relation with audiovisual labour and chains of rights. Historical node · Google DeepMind Read Veo 3 2025 Ideogram 3.0 updated model A new checkpoint only five weeks after 3.0 launch demonstrates continuous rather than annual model shipping. Historical node · Ideogram Read Ideogram 3.0 updated model 2025 HunyuanCustom Customized video generation combines multimodal reference inputs with the HunyuanVideo foundation. Historical node · Tencent Read HunyuanCustom 2025 Imagen 4 Imagen 4 is announced alongside Veo 3 and Lyria 2. The master sees this as evidence that image, video and audio generation are being organised as a collection of media capabilities within one platform, rather than remaining separate model categories with isolated interfaces and workflows. Historical node · Google DeepMind / Imagen team Read Imagen 4 2025 Flow A dedicated filmmaking interface unifies Veo, Imagen and Gemini into shot/scene-oriented creative workflows. Historical node · PRODUCT Read Flow 2025 HunyuanVideo-Avatar Audio-driven human animation extends open video generation into digital-human performance. Historical node · Tencent Read HunyuanVideo-Avatar 2025 FLUX.1 Kontext FLUX.1 Kontext extends text-to-image generation into a process that accepts both images and text, maintaining relationships among characters and scenes through successive edits. The master identifies a growing emphasis on reference-led and editing-centered models, beyond generation that begins only from a written prompt. Historical node · Black Forest Labs research team Read FLUX.1 Kontext 2025 Seedance 1.0 Native multi-shot storytelling, 1080p generation and fast inference push ByteDance into the frontier video tier. Historical node · ByteDance Seed Read Seedance 1.0 2025 V1 Video Model Midjourney expands from still images into image-to-video as a step toward interactive world models. V1 here identifies the video model, not the original V1 image model. Historical node · Midjourney Read V1 Video Model 2025 FLUX.1 Krea [dev] open weights An aesthetics-focused commercial model is distilled into an open FLUX-compatible checkpoint, connecting proprietary taste tuning to community infrastructure. Open weights do not imply unconditional authorisation for every use. Historical node · Krea / Black Forest Labs Read FLUX.1 Krea [dev] open weights 2025 Gemini 2.5 Flash Image Gemini 2.5 Flash Image develops the preservation of people and objects, combination of multiple images and conversational editing. It represents the growing importance of generation grounded in references during 2025: creative work increasingly involves continuing to work with existing visual subjects, rather than always producing an image from scratch. Historical node · Google DeepMind / Gemini team Read Gemini 2.5 Flash Image 2025 Qwen-Image A 20B image model emphasises Chinese/English text rendering and open-weight availability, contributing to the global open image ecosystem. This contributes multilingual capabilities to the global open-image-model ecosystem. Historical node · Alibaba Qwen Read Qwen-Image 2025 Lucid Origin Leonardo launches a new in-house model emphasising aesthetic diversity, prompt adherence and Full HD output. Historical node · Leonardo.Ai Read Lucid Origin 2025 Qwen-Image-Edit Semantic and appearance-preserving editing extends the base image model into precise textual editing and object transformation. Historical node · Alibaba Qwen Read Qwen-Image-Edit 2025 Realtime Video Video generation becomes faster-than-playback and interactive via canvas, webcam and screen-stream controls. Historical node · Krea Read Realtime Video 2025 HunyuanImage 2.1 A 17B open 2K text-to-image model combines MLLM text understanding, bilingual glyph-aware encoding and a refiner stage. Historical node · Tencent Read HunyuanImage 2.1 2025 HunyuanImage 3.0 Tencent shifts toward native multimodal image generation with open weights and stronger instruction understanding. Historical node · Tencent Read HunyuanImage 3.0 2025 Nano Banana expands to Search / NotebookLM / Photos A specialist image model becomes infrastructure inside search, knowledge tools and consumer photo products. Historical node · Google Read Nano Banana expands to Search / NotebookLM / Photos 2025 HunyuanVideo 1.5 A lighter 8.3B model targets consumer-grade GPUs, showing open video generation moving from research-scale hardware toward local use. Consult the project documentation for actual deployment requirements. Historical node · Tencent Read HunyuanVideo 1.5 2026 Niji 7 Anime generation gains stronger coherence, fine line work and repeatable characters while remaining a distinct specialist model line. Historical node · Midjourney / Spellbrush Read Niji 7 2026 HunyuanImage 3.0 Instruct / Instruct-Distil Native image generation gains instruction reasoning, image-to-image editing and a distilled deployment variant. Historical node · Tencent Read HunyuanImage 3.0 Instruct / Instruct-Distil 2026 Seedance 2.0 official launch Seedance 2.0 uses text, images, audio and video as reference inputs and combines multi-shot generation with editing controls. Its official launch announcement is dated 12 February 2026, correcting the June date in the research master. Historical node · ByteDance Seed Read Seedance 2.0 official launch 2026 Firefly Custom Models public beta Adobe productizes custom style/character/photographic models trained on a creator’s own images inside Firefly. Historical node · Adobe Read Firefly Custom Models public beta 2026 ChatGPT Images 2.0; GPT-Image-2 API The master places ChatGPT Images 2.0 and GPT-Image-2 within a move toward production workflows, emphasising instruction following, layouts, text and higher-resolution output. It also records availability through developer and coding interfaces. The official announcement confirms the release date; the wider interpretation and individual capability claims require separate assessment. Historical node · OpenAI image generation team Read ChatGPT Images 2.0; GPT-Image-2 API 2026 GPT-Image-2 / ChatGPT Images 2.0 The master records GPT-Image-2's introduction through the API and Codex, with an emphasis on complex visual tasks, layouts, text, instruction following and output up to 2K. Its interpretation is a shift from producing an attractive image toward generating assets for products and design systems. Current model documentation is linked separately from the historical account. Historical node · OpenAI image generation team Read GPT-Image-2 / ChatGPT Images 2.0 2026 Muse Image launch & Muse Video preview Meta releases Muse Image with search and code tools and previews a related video model with native audio. The released image model and the early video preview have different availability states; they should not be treated as a simultaneous public release of both products. Historical node · Meta Superintelligence Labs Read Muse Image launch & Muse Video preview 2026 Hailuo 3.0 available via Runway API Runway adds third-party Hailuo 3.0 video generation to its developer platform, illustrating how creative platforms also route models from other suppliers. This node dates API availability on Runway, not the model’s original release. Historical node · Runway Dev Read Hailuo 3.0 available via Runway API 2026 ChatGPT Images 2.5 / GPT-Image-2.5 Flare & Sunburst The announcement presents improvements in detail, subject preservation, multi-turn editing and speed, alongside Sketch, Templates and comments placed directly on images. It introduces Flare and Sunburst API variants. The master reads this as a convergence of generation, reference, annotation, editing, templates and collaboration within the image-making interface of 2026. Historical node · OpenAI image generation team Read ChatGPT Images 2.5 / GPT-Image-2.5 Flare & Sunburst