Home › Blog › Image & Video AI Models
🎨

Image & Video AI Models:
The Complete Guide

Every model, benchmark, player, startup, architecture, dataset, monetization model, and research frontier — for generative image and video AI in 2026.

FL
FrontierAGI Team
Generative AI Computer Vision Deep Dive

Generative image and video AI has moved from research curiosity to industrial infrastructure in less than four years. Stable Diffusion made image synthesis democratised in 2022. By 2026, you can generate a photorealistic 4K video clip from a text prompt in under 60 seconds — and companies are spending billions racing to make that generation faster, longer, more controllable, and natively multimodal. This article maps the entire landscape: every major model, every key player, the architectures that power them, how to build one, how to monetize one, and where research is headed next.

🕰️ A Decade of Visual AI — Complete History

From GANs that could barely produce a recognisable face to real-time 4K video generation in twelve years. Here's every landmark moment, in order:

Image Model
Video Model
Architecture / Research
Open Source
🎭 Era 1 — GAN Era (2014–2020)
2014
Architecture Jun 2014
GANs — Generative Adversarial Networks
Ian Goodfellow et al. · NeurIPS 2014
BreakthroughTwo networks (generator + discriminator) in adversarial training — the first viable deep generative framework.
ImpactLaunched the entire field of learned image synthesis. Every major generative image model for the next 8 years built on or competed with GANs.
ArchitectureFoundational
2015
Image Jun 2015
Deep Dream
Google Brain · Feature visualisation goes viral
BreakthroughInceptionism / gradient ascent on CNN activations produces hallucinatory imagery — first mainstream AI art moment.
ImpactIntroduced the public to the idea that neural networks "see" the world. Cultural turning point for AI art.
ImageCultural Milestone
2016
Architecture Nov 2016
pix2pix
UC Berkeley · Image-to-image translation with cGANs
BreakthroughConditional GAN learns to translate between any paired image domain (sketch→photo, day→night, map→aerial).
ImpactTemplate for all conditional image generation. Spawned CycleGAN, Vid2Vid, and the entire image editing-with-AI lineage.
ArchitectureImage Editing
2017
Image Oct 2017
Progressive GAN
NVIDIA · High-resolution face synthesis via staged training
BreakthroughTrains GAN from 4×4 up to 1024×1024 by progressively adding resolution layers — first photorealistic face synthesis at HD.
Impact1024×1024 generated faces indistinguishable from photos. Set off the deepfake conversation. Direct ancestor of StyleGAN.
ImageNVIDIA
2018
Image Sep 2018
BigGAN
DeepMind · Class-conditional ImageNet at 512px
BreakthroughLargest GAN trained to date — 512px class-conditional generation across 1000 ImageNet categories at unprecedented quality.
ImpactDemonstrated scaling laws apply to GANs. Introduced truncation trick for quality/diversity tradeoff.
ImageDeepMind
2019
Image Feb 2019
StyleGAN
NVIDIA · Style-based generator with mapping network
BreakthroughIntroduces W-space latent mapping and AdaIN style injection — fine-grained control over coarse (pose) and fine (texture) attributes separately.
Impact"This person does not exist" goes viral. StyleGAN2 and StyleGAN3 remain best-in-class for face synthesis. Inspires latent editing research (GANSpace, InterFaceGAN).
ImageNVIDIAFaces
2020
Architecture Jun 2020
DDPM — Denoising Diffusion Probabilistic Models
Ho et al. · NeurIPS 2020 · The architecture that changed everything
BreakthroughLearns to reverse a noise-addition process step by step. Achieves image quality surpassing GANs without adversarial training instability.
ImpactThe foundation of every modern image and video generation model. Stable Diffusion, DALL-E 2, Imagen, Sora — all descendants.
ArchitectureFoundationalMilestone
🌅 Era 2 — Diffusion Dawn (2021–2022)
2021
Architecture Jan 2021
CLIP
OpenAI · Contrastive Language-Image Pretraining
BreakthroughTrained on 400M image-text pairs to align visual and language embeddings. Zero-shot image classification rivalling supervised models.
ImpactUnlocked text-guided image generation. The "steering wheel" plugged into diffusion models to make them prompt-responsive. Powers FID evaluation and aesthetic scoring to this day.
ArchitectureOpenAIMultimodal
Image Jan 2021
DALL-E 1
OpenAI · First public text-to-image model
BreakthroughUses a dVAE to tokenise images, then a GPT-style transformer to generate image tokens from text. 256px, surprising prompt adherence.
ImpactFirst model that could reliably generate "an armchair in the shape of an avocado." Proved text-to-image was viable. Named the product category.
ImageOpenAIText-to-Image
Image Dec 2021
GLIDE
OpenAI · Diffusion + CLIP guidance — photorealism breakthrough
BreakthroughCombines diffusion models with classifier-free guidance and CLIP text conditioning. First model to produce photorealistic scenes from prompts.
ImpactOutperformed DALL-E 1 dramatically. Human evaluators preferred it 3:1. The proof-of-concept that led directly to DALL-E 2.
ImageOpenAIDiffusion
2022
Image Apr 2022
DALL-E 2
OpenAI · CLIP-guided diffusion at 1024px
BreakthroughCLIP image embeddings as conditioning signal. Supports inpainting, outpainting, and image variations. 4× higher resolution than DALL-E 1.
ImpactFirst model to create an industry — spawned Midjourney, Stability AI, and dozens of startups. Mainstream media moment: "AI can draw anything."
ImageOpenAI1024px
Image May 2022
Imagen
Google Brain · Cascaded diffusion + large language model encoder
BreakthroughUses T5-XXL (frozen) as text encoder — much stronger semantic understanding than CLIP. Cascaded diffusion pipelines for high resolution.
ImpactBeat DALL-E 2 on FID on COCO. Introduced DrawBench benchmark. Demonstrated LLM text encoders are superior for complex prompts.
ImageGooglePhotorealism
Open Source Aug 2022
Stable Diffusion 1.x ⭐
CompVis / Stability AI · Latent Diffusion open-sourced
BreakthroughLatent Diffusion Model (LDM) compresses images to 8× smaller latent space before denoising — runs on a consumer GPU for the first time. Full weights released publicly.
ImpactThe most consequential open-source AI release ever in the visual domain. Sparked: Civitai (100K+ community models), DreamBooth, ControlNet, A1111 WebUI with 50M downloads, and a Cambrian explosion of visual AI applications.
Open SourceMilestoneDemocratisation
Image Oct 2022
Midjourney v3/v4
Midjourney · Aesthetic quality leap via Discord
BreakthroughOptimises for human aesthetic preference rather than raw FID. Discord-native UX creates massive community feedback loop.
ImpactBecame the tool of choice for concept artists, designers, and creators. 180M+ users by 2026. $200M ARR with zero VC funding.
ImageMidjourneyAesthetic
Video Oct 2022
Make-A-Video
Meta AI · First text-to-video diffusion model
BreakthroughExtends a text-to-image diffusion model into the temporal dimension without video-text pairs — leverages image-text supervision only.
ImpactFirst demonstration that video generation from text was tractable at all. 16 frames, 256px — primitive but proved the concept.
VideoMetaFirst
🏭 Era 3 — Commercialisation (2023)
2023
Architecture Feb 2023
ControlNet
Lvmin Zhang · Structure-guided diffusion conditioning
BreakthroughAdds trainable copy of SD encoder to accept structural conditioning: pose skeletons, depth maps, Canny edges, scribbles. Any structure can guide generation.
ImpactMade AI image generation useful for professional workflows. Designers could now maintain composition while changing style. GitHub's #1 trending repo in 2023.
ArchitectureControllabilityOpen Source
Open Source Apr 2023
Segment Anything Model (SAM)
Meta AI · Universal image segmentation at zero-shot
BreakthroughTrained on 1B masks — segments any object in any image with a click, box, or text prompt. Zero-shot generalisation across domains.
ImpactBecame the standard preprocessing layer for editing pipelines. SAM 2 (2024) extends to video. Used in robotics, medical imaging, and autonomous driving.
Open SourceMetaSegmentation
Open Source Jul 2023
SDXL 1.0
Stability AI · 1024px dual-encoder latent diffusion
Breakthrough3× larger U-Net, dual text encoders (CLIP-L + OpenCLIP-G), 1024px native resolution, refiner model for detail. Open weights.
ImpactSet the open-source image quality bar for 18 months. Ecosystem of SDXL LoRAs and checkpoints on Civitai exceeded 50K models.
Open SourceStability AI1024px
Video Jun 2023
Runway Gen-2
Runway · First commercially viable text/image-to-video
BreakthroughMulti-modal conditioning (text, image, or video input). Consistent motion and scene structure for 4-second clips. First product deployed in real film productions.
ImpactValidated the professional video generation market. Used in multiple award-winning short films. Triggered $236M Series D for Runway.
VideoRunwayCommercial
Image Oct 2023
DALL-E 3
OpenAI · Complex prompt adherence + ChatGPT integration
BreakthroughRe-captioned training data with GPT-4V — dramatically improved prompt following. Integrated directly into ChatGPT. Text rendering 10× better than previous generation.
ImpactBundled with ChatGPT gave DALL-E 3 distribution to 100M+ users overnight. Set the standard for prompt adherence that competitors still chase.
ImageOpenAIText Rendering
Video Nov 2023
Pika Labs 1.0
Pika · Creator-focused video generation goes viral
BreakthroughDiscord-native video generation with scene-modify, expand, and fill controls. Accessible quality in seconds. $55M raised at launch.
ImpactBrought video generation to 500K+ creators within weeks. Validated consumer video gen market. Triggered Runway Gen-3 and Kling acceleration.
VideoStartupCreator Economy
Open Source Nov 2023
Stable Video Diffusion (SVD)
Stability AI · First open-weight video generation model
BreakthroughFine-tunes SDXL image encoder into temporal domain. Generates 14 or 25 frames at 576×1024. Full open weights released.
ImpactFirst open-source video model usable locally. Triggered CogVideoX, Open-Sora, and the open video research ecosystem.
Open SourceVideoStability AI
⚔️ Era 4 — The Video Wars (2024)
2024
Video Feb 2024
Sora ⭐
OpenAI · 60-second 1080p cinematic video — industry shock moment
BreakthroughVideo Diffusion Transformer (DiT) operating on spatial-temporal patches. Understands object permanence, physics, scene consistency across 60 seconds.
ImpactShocked Hollywood, VFX studios, and the entire creative industry. "The iPhone moment for AI video." Triggered emergency strategy reviews at every major studio. Launched the global video AI arms race.
VideoOpenAIMilestoneDiT
Open Source Mar 2024
Stable Diffusion 3
Stability AI · Multimodal Diffusion Transformer (MMDiT)
BreakthroughReplaces U-Net with Multimodal DiT. Separate streams for image and text tokens with bidirectional attention between them. 3B parameters. Excellent text rendering in images.
ImpactProved DiT architecture superiority over U-Net at scale. Open weights set new open-source benchmark. Template for SD 3.5 and community fine-tuning.
Open SourceDiTStability AI
Image Aug 2024
FLUX.1
Black Forest Labs · Rectified Flow Transformer — new state of the art
BreakthroughRectified Flow Matching + Transformer architecture from original SD creators (Robin Rombach et al.). Dramatically better anatomy, hands, and text rendering. Three tiers: [pro], [dev] (open weights), [schnell] (distilled, 4 steps).
ImpactImmediately became community favourite. [dev] weights downloaded 2M+ times in first month. Set new standard for open image models that SD 3.5 had to match.
ImageOpen SourceFlow Matching
Video Jun 2024
Kling 1.0
Kuaishou · China's Sora challenger — 30 sec at 1080p
Breakthrough3D Variational Autoencoder for spatial-temporal consistency. 30 seconds at 1080p with comparable motion quality to Sora. Strong free tier drives viral adoption.
ImpactFirst Chinese model to credibly compete with Western frontier video models. Forced OpenAI and Runway to accelerate releases. 10M users in 3 months.
VideoChinaKuaishou
Open Source Aug 2024
CogVideoX-5B
THUDM (Tsinghua) · Best open-weight video model of 2024
Breakthrough5B parameter video DiT with expert transformer blocks. Full open weights including 5B and 2B variants. Trained on curated OpenVid-1M dataset.
ImpactFirst open video model with quality approaching commercial offerings. Enabled academic video research. Spawned CogVideoX fine-tuning ecosystem.
Open SourceVideoDiT
Video Oct 2024
Movie Gen
Meta AI · Joint video + synchronized audio generation
Breakthrough30B parameter model that jointly generates 16-second HD video with temporally aligned audio. Fine-grained editing of real videos. Open research paper, no public weights.
ImpactFirst demonstration of true audio-video co-generation. Sets direction for 2026's multimodal video systems. Research baseline for the industry.
VideoMetaAudio
Video Dec 2024
Veo 2
Google DeepMind · 4K · 120 seconds · Physics-aware
Breakthrough4K resolution, 120-second clips, physics simulation, cinematic camera controls. Highest VBench scores of any model at launch. Integrated into Vertex AI and VideoFX.
ImpactDethroned Sora on benchmark leaderboards. Demonstrated Google's TPU infrastructure advantage for video at scale. Available to enterprise customers via API.
VideoGoogle4KPhysics
🌐 Era 5 — Convergence & Realtime (2025–2026)
2025
Image Feb 2025
GPT-4o Native Image Generation
OpenAI · Autoregressive + diffusion hybrid; conversational editing
BreakthroughAutoregressive token generation (not diffusion) for images within the same model as text. Conversational image editing ("make the hat blue"), OCR and text retention across edits.
ImpactCategory-defining: image editing as dialogue. 1M+ images generated in first 24 hours. Made accurate text-in-image trivially easy. Disrupted standalone image generation tools.
ImageOpenAIEditingAutoregressive
Video Mar 2025
Runway Gen-4
Runway · Multi-shot consistency + face transfer at 4K
BreakthroughConsistent character identity across multiple shots and camera angles. Act-One: facial performance transfer from actor reference. 4K output. Professional film-grade controls.
ImpactFirst video model to support multi-shot productions — the missing piece for professional film use. Used in Tribeca Film Festival selections. Industry adoption accelerates.
VideoRunwayProfessional
Image Jun 2025
Ideogram 2.5
Ideogram · World-best text rendering + typography control
BreakthroughFine-grained control over font, weight, position, and styling of text within images. Accurate multi-line text rendering — previously the hardest problem in image generation.
ImpactMade AI practical for graphic design, marketing, and typographic work. Preferred over DALL-E 3 and FLUX for any text-heavy creative work.
ImageTypographyDesign
Image Sep 2025
Midjourney v7
Midjourney · Personalisation system + 4096px native
BreakthroughPersonalisation trains a user-specific style model from their rating history. 4096px native. Style reference locking for brand consistency across prompts.
ImpactFirst model with true per-user aesthetic adaptation. Enterprise brands can "own" a visual style. Continues Midjourney's dominance of the aesthetic image market.
ImageMidjourneyPersonalisation
2026
Open Source Early 2026
SD 3.5 + Flow Matching Era
Stability AI / Community · Flow matching becomes the open standard
BreakthroughFlow matching (Rectified Flow) replaces DDPM as the default training paradigm for new open image models. 8-step inference at quality previously requiring 50 steps.
ImpactConsumer GPU generation under 2 seconds. Open fine-tuning ecosystem catches up to commercial quality in most domains. Inference cost drops 80% vs 2022 SD.
Open SourceFlow MatchingEfficiency
Video Mid 2026
Real-Time Video Generation Era
Industry-wide · Sub-10s video becomes commercially standard
BreakthroughDistillation + flow matching brings 720p 5-second video generation under 10 seconds on H100. Multiple providers offer real-time preview modes.
ImpactInteractive video editing becomes viable. Game asset generation, social media creation, and advertising workflows transformed. Video generation as ubiquitous as image generation was in 2023.
VideoReal-Time2026

📊 The Market in Numbers

$12B
Generative Image/Video AI market (2026)
$48B
Projected market by 2030 (CAGR ~40%)
3B+
Images generated per day globally
180M+
Midjourney registered users
$6B
VC investment in generative media (2024–2025)
75%
Marketers using AI image tools by Q1 2026

Three macro forces are driving the expansion: (1) compute democratisation — inference costs for image generation have dropped ~10× since 2022; (2) quality-accuracy convergence — the gap between "AI-generated" and "real" is closing at the pixel level; and (3) workflow integration — tools are moving from standalone apps into the creative software stack (Adobe, Figma, Canva, DaVinci Resolve).

🖼️ Image Generation
Single-frame, spatial only
512×512 to 16K native resolution
1–5 seconds per image (2026)
~5–20 GB model weights
Mature benchmark ecosystem (FID, CLIP, HPSv2)
High inference efficiency on consumer GPUs
🎬 Video Generation
Temporal consistency adds a 3rd dimension
720p to 4K, 5–120 seconds
30 seconds – 10 minutes generation time
~50–200 GB model weights
Benchmark ecosystem still maturing (VBench, EvalCrafter)
Requires A100/H100 class GPUs

🖼️ Image Generation & Editing Models

The image model landscape has bifurcated into two tiers: frontier commercial models with proprietary training data and RLHF-tuned quality, and a thriving open-source ecosystem led by the Stable Diffusion family. Here is the complete 2026 catalogue:

Model Company Type Arch Max Res Access Notable Strength
DALL-E 3 OpenAI Image Gen Diffusion 1792×1024 API / ChatGPT Prompt adherence, text rendering
GPT-4o Image OpenAI Image Gen Edit Autoregressive + Diffusion 1024×1024 ChatGPT / API Conversational editing, OCR retention
Imagen 3 Google DeepMind Image Gen Cascaded Diffusion 2048×2048 Gemini / Vertex AI Photorealism, detail fidelity
Midjourney v7 Midjourney Image Gen Diffusion (proprietary) 4096×4096 Discord / Web Aesthetic quality, stylisation
Adobe Firefly 3 Adobe Image Gen Edit Diffusion (fine-tuned) 4096×4096 Photoshop / API Commercial safe, generative fill
Stable Diffusion 3.5 Stability AI Image Gen Flow Matching (DiT) Flexible Open Source Local deployment, fine-tuning
FLUX.1 [pro] Black Forest Labs Image Gen Rectified Flow Transformer Flexible API / OSS (dev) Typography, hands, anatomical accuracy
Ideogram 2.5 Ideogram Image Gen Diffusion 2048×2048 API / Web Best-in-class text rendering
Recraft V3 Recraft Image Gen Edit Diffusion 2048×2048 API / Web SVG gen, brand consistency, style locks
Gemini Imagen (Whisk) Google Image Gen Cascaded Diffusion 1024×1024 Google Labs Subject-style-scene remixing
Meta Emu Edit 2 Meta AI Edit Diffusion 1024×1024 Research / API Instruction-guided editing, style transfer
Kandinsky 3 Sber AI Image Gen Latent Diffusion 1024×1024 Open Source Russian-language prompt support
PixArt-Σ PixArt Research Image Gen DiT 4K Open Source Efficient training, high resolution
InstructPix2Pix UC Berkeley Edit Diffusion (instruction) 512×512 Open Source Text-guided image editing baseline

🎬 Video Generation & Editing Models

Video generation is the 2025–2026 battleground. The challenge isn't just generating frames — it's generating coherent motion across time while keeping identity, physics, and scene context stable. Here is the 2026 video model landscape:

Model Company Type Arch Max Length Access Notable Strength
Sora OpenAI Video Gen Video Diffusion Transformer (DiT) 60 sec / 1080p ChatGPT Pro / API Cinematic motion, scene transitions
Veo 2 Google DeepMind Video Gen Cascaded Video Diffusion 120 sec / 4K Vertex AI / VideoFX Physics simulation, 4K quality
Runway Gen-4 Runway Video Gen Edit Latent Video Diffusion 30 sec / 4K Subscription / API Multi-shot consistency, professional controls
Kling 1.6 Kuaishou Video Gen 3D Variational Autoencoder 30 sec / 1080p API / Web Motion quality, free tier availability
Pika 2.2 Pika Labs Video Gen Edit Diffusion 15 sec Web / API Pikaffects, scene modification, fast iteration
Luma Dream Machine 2 Luma AI Video Gen Diffusion Transformer 30 sec / 1080p Web / API Consistent character motion, camera moves
Hailuo (MiniMax) MiniMax Video Gen Video DiT 30 sec / 1080p Web / API Realistic human motion
CogVideoX-5B THUDM (Tsinghua) Video Gen DiT 10 sec Open Source Best open-weight video model as of 2025
Open-Sora 2.0 HPC-AI Tech Video Gen Spatial-temporal DiT 60 sec Open Source Full open pipeline, research baseline
Wan (Alibaba) Alibaba Video Gen Video Flow Matching 10 sec / 720p Open Source Efficient open video model
InVideo AI InVideo Edit Pipeline (LLM + diffusion) Unlimited scenes Web subscription Script-to-video, voiceover, stock integration
Descript Underlord Descript Edit Pipeline Unlimited App subscription Transcript-based editing, AI overdub

📐 Benchmarks: How Models Are Evaluated

Benchmarking generative visual AI is fundamentally harder than benchmarking language models. You can't ask "is this answer correct?" — you're evaluating perceptual quality, semantic alignment, aesthetic coherence, and temporal consistency. Multiple complementary metrics are required.

Image Quality Benchmarks
HPSv2 Score
FLUX
FID (lower=better)
Imagen
CLIP Score
SD3.5
GenEval
DALL-E 3
T2I-Compbench
Ideogram
Video Quality Benchmarks
VBench Overall
Veo 2
EvalCrafter
Sora
Temporal Consist.
Runway
Motion Smoothness
Kling
Human Preference
Pika
Key Metrics Explained

FID (Fréchet Inception Distance) — Distribution distance between real and generated image features. Lower = more realistic. Standard reference since 2017.

CLIP Score — Cosine similarity between image and prompt in CLIP embedding space. Measures prompt adherence.

HPSv2 — Human Preference Score trained on 798K image-preference pairs. Best proxy for "what humans actually like."

VBench — 16-dimension video quality benchmark: subject consistency, background, motion, aesthetic, temporal flicker, scene.

EvalCrafter — 700 prompts × 4 criteria (visual quality, text alignment, motion, action semantics).

🏢 Major Players

OpenAI
Image + Video
DALL-E 3 GPT-4o Image Sora
Bundled distribution through ChatGPT (180M+ users) gives OpenAI structural advantage. Sora's cinematic quality is unmatched but speed lags behind Runway. GPT-4o's conversational image editing is a category-creator.
Google DeepMind
Image + Video
Imagen 3 Veo 2 VideoFX Whisk
Veo 2 leads on 4K video benchmarks. Imagen 3 integrated into Gemini Advanced. Google's TPU infrastructure gives cost advantage at scale. Deep Search integration adds multimodal grounding.
Midjourney
Image
v6.1 v7 Niji 7
$200M ARR, 180M users, no outside funding. Dominates creative/aesthetic image generation. v7 introduces "Personalization" for fine-tuned style locks. Strong Discord community creates self-reinforcing flywheel.
Adobe
Image + Video (Editing Focus)
Firefly 3 Firefly Video Generative Fill
Only major player trained exclusively on licensed/owned content — crucial for commercial rights. Deep Photoshop/Premiere/After Effects integration drives enterprise adoption. $100/mo Creative Cloud includes unlimited Firefly credits.
Stability AI
Image (Open Source Leader)
SD 3.5 SDXL SD Video
Pioneer of open-source image generation. SD 3.5 uses flow matching (DiT) architecture. Company went through restructuring in 2024 but model releases continue. Huge ecosystem of fine-tuned variants (Civitai has 100K+ models built on SD).
Black Forest Labs
Image
FLUX.1 [pro] FLUX.1 [dev] FLUX.1 [schnell]
Spun out of Stability AI by SD's original creators (Robin Rombach et al.). FLUX uses Rectified Flow Transformer — superior anatomy, text, and hand generation. [dev] and [schnell] are open weights. $100M raised in 2024.
Runway
Video
Gen-3 Alpha Gen-4 Act-One
Industry standard for professional video generation. Gen-4 introduces multi-shot consistent characters. Act-One enables facial performance transfer. $236M raised, Film Academy partnerships. Targets professional film/TV workflows.
Meta AI
Image + Video (Research)
Emu Edit 2 Movie Gen Segment Anything 2
Movie Gen (Oct 2024) generates 16-second HD video with synchronized audio. SAM 2 enables video segmentation at scale. Meta releases research weights — critical infrastructure for the open ecosystem.
Kuaishou (快手)
Video
Kling 1.0 Kling 1.6
China's answer to Sora. Kling 1.6 achieves competitive motion quality at significantly lower inference cost. Strong free tier drives viral adoption. 3D VAE architecture for spatial-temporal consistency.
ElevenLabs
Audio + Video Synthesis
Voice Clone Sound Effects Video Dubbing
While primarily audio, ElevenLabs' lip-sync and video dubbing tools are core to full video production pipelines. Integrates with Runway and Pika for end-to-end generation.

🚀 Major Startups & Funding Activity

Over $6B has flowed into generative image and video AI startups since 2023. Here are the most significant companies and rounds:

Pika Labs
$80M
Series B · 2024 · Valuation ~$470M
Andreessen Horowitz, Khosla Ventures
Runway
$236M
Series D · 2024 · Valuation ~$1.5B
General Atlantic, Google, Salesforce Ventures
Black Forest Labs
$100M
Seed/Series A · 2024
General Catalyst, A16Z
Ideogram
$80M
Series B · 2024 · Valuation ~$700M
Andreessen Horowitz, Index Ventures
Luma AI
$43M
Series B · 2024
Andreessen Horowitz, Kleiner Perkins
Stability AI
$101M
Seed + Strategic · 2022
Coatue, Lightspeed, O'Shaughnessy
HeyGen
$60M
Series A · 2024 · Valuation ~$500M
Benchmark, 8VC
Synthesia
$90M
Series C · 2024 · Valuation ~$1B
Accel, NVentures (NVIDIA)
Recraft
$12M
Series A · 2024
Khosla Ventures
Higgsfield AI
$8M
Seed · 2024
Andreessen Horowitz
Krea AI
$19M
Series A · 2024
Index Ventures, Abstract
Captions
$60M
Series B · 2024 · Valuation ~$500M
Kleiner Perkins, Sequoia

Investor thesis patterns: a16z dominates early stage ($100K–$80M), Kleiner Perkins and Index focus on Series B+ with proven revenue. NVIDIA's NVentures is strategic investor in 12+ companies (infrastructure moat). Google and Microsoft invest in ecosystem plays that funnel compute revenue. Consolidation accelerating — Adobe acquired Firefly team, NVIDIA acquired Eleven Labs-adjacent companies.

🗂️ Data Sources & Training Datasets

The quality of generative models is inseparable from the quality and scale of their training data. Here's a complete map of datasets used to build image and video AI:

Dataset Scale Type License Used By Notes
LAION-5B 5.85B image-text pairs Image + Caption CC-BY 4.0 SD 1.x, SD 2.x, DALL-E 2 (partial) Scraped from Common Crawl. Safety concerns raised in 2023 (CSAM). Filtered LAION-5B-cleaned more commonly used.
LAION Aesthetics 600M images High-aesthetic image-text CC-BY 4.0 SD 2.x, SDXL Filtered from LAION-5B using CLIP aesthetic score ≥ 4.5. Key for quality improvement.
COYO-700M 700M image-text pairs Image + Caption Apache 2.0 FLUX, SD 3.x Better filtered than LAION. NSFW filtered version widely used. Korean alternative to LAION.
DataComp-1B 1.28B samples Image + Caption CC-BY 4.0 OpenCLIP, research UNC / UW benchmark dataset for understanding data curation choices.
JFT-3B (Google) 3B images Image + Labels Proprietary Imagen series Internal Google dataset. Not public. Key competitive advantage for Imagen quality.
Adobe Stock ~400M licensed assets Image + Video + Vectors Proprietary (licensed) Firefly Only large-scale commercially licensed training data. Key differentiator for enterprise use.
WebVid-10M 10M video-text pairs Video + Caption CC-BY 4.0 Early video models Scraped from stock video sites. Superseded by larger private datasets for frontier models.
HD-VILA-100M 100M video clips Video + Caption Research only Research video models Microsoft Research. High-definition variety for video understanding and generation.
OpenVid-1M 1M high-quality clips Video + Text CC-BY 4.0 CogVideoX, Open-Sora Quality-filtered, aesthetically scored. Open alternative to proprietary video datasets.
Internal Curation (Sora, Veo, Kling) Unknown (billions est.) Video + Caption Proprietary Sora, Veo, Kling All frontier video models train on proprietary datasets curated from the internet + licensed content. This is the core competitive moat.

The data problem is now the primary bottleneck for video generation. While language models can train on essentially all public text, video data is: (1) much larger per sample (~1GB per minute of 4K), (2) unevenly captioned, (3) temporally complex to annotate, and (4) increasingly locked behind platform rights. The companies that curate better video datasets — with better scene/motion/caption alignment — will win the next generation of video models.

⚙️ Model Architectures

Four architectures power generative image and video AI. Understanding them helps you choose the right approach for your use case and budget.

1. Denoising Diffusion Probabilistic Models (DDPMs)

The dominant architecture since 2022. A diffusion model learns to reverse a noise-addition process: given a completely noisy image, it predicts what the clean image should look like — step by step.

Pure Noise
←
t=800
←
t=500
←
t=200
←
Clean Image

Denoising process: model iteratively removes noise guided by text conditioning

Latent Diffusion Models (LDMs) — the key innovation in Stable Diffusion — perform denoising in a compressed latent space rather than pixel space, making generation 10–100× faster. The encoder compresses a 512×512 image to a 64×64 latent; the denoiser operates on that; the decoder expands back to pixels.

2. Diffusion Transformers (DiT)

Replaces the U-Net denoiser with a pure Transformer architecture. Introduced by Peebles & Xie (2022), DiT scales more predictably with compute (like LLMs) and handles larger contexts. Used by: SD 3.x, FLUX, Sora, CogVideoX, PixArt. This is the dominant 2025–2026 architecture.

📝
Text Encoder
T5-XXL / CLIP
🗜️
VAE Encoder
Image → Latent
🔀
DiT Blocks
Multi-Head Attn
🔊
Noise Scheduler
DDIM / Flow
🖼️
VAE Decoder
Latent → Pixels

3. Flow Matching

A cleaner mathematical alternative to DDPM. Instead of learning to reverse a noisy process, the model learns a velocity field that flows noise smoothly toward data. Advantages: fewer inference steps (4–8 vs 20–50 for DDPM), straighter trajectories, more stable training. Used by: FLUX.1, SD 3.x, Wan (Alibaba). Becoming the default for new image models.

4. Autoregressive Transformers

Generate images token-by-token like text. DALL-E 1 (2021) used this approach; it was largely replaced by diffusion models for quality reasons. Now returning at scale: GPT-4o's image generation uses a hybrid autoregressive + diffusion approach. Advantages: natural integration with language models, supports interleaved image-text generation. Disadvantages: slower, less fine-tuned aesthetic control.

Architecture Comparison
DDPM (U-Net)
Mature
DiT
Best
Flow Matching
Rising
Autoregressive
Niche
GAN (legacy)
Fading

🔨 How to Build an Image/Video AI Model From Scratch

Building a frontier generative model requires months of engineering and tens of millions in compute. But understanding the pipeline is essential for anyone building in this space — whether you're fine-tuning, building applications, or doing research.

1. Define Your Scope and Architecture

Image-only or video? Fine-tune existing weights or train from scratch? For most startups, fine-tuning SD 3.5 or FLUX.1 [dev] on domain-specific data is far more practical than pre-training. Pre-training a competitive model from scratch costs $2–50M in compute. Choose: DiT for scalability, U-Net for compatibility with existing LoRA ecosystem, flow matching for efficiency.

2. Assemble and Clean Your Dataset

Scrape or license images/videos with captions. Run NSFW filters (CLIP-based), aesthetic scorers (LAION aesthetic predictor), de-duplication (minhash LSH), and resolution filters. For video: extract scene transitions, compute optical flow to identify static clips. Aim for quality over quantity — 1M well-captioned images beat 10M poorly-captioned ones.

3. Caption Everything with a VLM

Re-caption your dataset using a vision-language model (LLaVA, InternVL, CogVLM) to generate dense, accurate captions. This step dramatically improves text-alignment. SD 3.5 and FLUX both used synthetic re-captioning as a key training improvement. Budget: ~$0.001 per image using API-based captioning, or run locally on 8× A100.

4. Train the Text Encoder

Use T5-XXL (4.7B params) for semantic understanding and CLIP ViT-L for visual alignment. For frontier models, freeze these or fine-tune minimally — the visual quality bottleneck is in the denoiser/generator, not the text encoder. Most startups simply reuse OpenCLIP weights.

5. Train the VAE (or reuse)

The VAE compresses images into latent space. SDXL's VAE (4-channel) is widely reused. SD 3.x upgraded to 16-channel VAE for better detail preservation. Training a VAE from scratch on 50M images takes ~1 week on 64 A100s. For most projects: reuse existing VAE, only pre-train if domain is highly domain-specific (medical, satellite).

6. Pre-train the Denoiser/Generator

This is the main model. For a DiT: stack transformer blocks with adaptive layer norm conditioning (adaLN). Scale: 0.6B–8B parameters. Training: mixed-precision (BF16), gradient checkpointing, flash attention. Estimate: $2–5M compute for a 1B parameter model on 1B images. Use cosine noise schedule for DDPMs, or ODE solvers for flow matching.

7. RLHF / DPO Alignment

Post-train with human preference data. Collect pairwise comparisons (A vs B for the same prompt). Train a reward model on 50K–500K pairs. Fine-tune the generator with DPO (Direct Preference Optimization) or DDPO (for diffusion). This step is responsible for ~30–40% of perceptual quality improvement on human benchmarks.

8. Safety Filtering and Evaluation

Train/deploy a safety classifier to block NSFW output, harmful prompts, and intellectual property violations. Evaluate on: FID, CLIP score, HPSv2, T2I-Compbench. Human eval via A/B tests on 500+ prompts across 12 categories (people, animals, objects, text, scenes, abstract). Iterate on both safety and quality. Red-team for adversarial prompts.

9. Optimize Inference

Quantize to INT8/INT4 (bitsandbytes, AWQ). Apply model distillation to reduce inference steps (SDXL-Turbo: 4-step distillation vs 50-step original). Use FlashAttention-3 for 2× throughput. Cache attention KV for similar prompts. Deploy on NVIDIA TensorRT for 3–8× latency reduction vs naive PyTorch. Target: <2 seconds for 1024×1024 on A10G.

10. Fine-tuning Ecosystem (LoRA, DreamBooth, IP-Adapter)

Your users will want customization. Implement LoRA (Low-Rank Adaptation) support — 4–16 rank adapters that let users fine-tune specific styles on 20–100 images. DreamBooth for subject-specific fine-tuning. IP-Adapter for image-conditioned generation. ControlNet for structural conditioning (pose, depth, edge). This adapter ecosystem is what made SD dominant — replicate it.

How Stable Diffusion Works
Explained from first principles · Computerphile
Sora: How OpenAI's Video AI Works
Architecture breakdown · Yannic Kilcher
Diffusion Models Explained
Score matching & DDPM · AssemblyAI
FLUX.1 Architecture Deep Dive
Rectified flow transformers · AI Coffee Break

💰 Monetization Models

🔢
Credit / Token System
Users buy bundles of credits; each generation consumes credits based on resolution and complexity. Predictable revenue, low entry barrier.
Midjourney: $10/mo = ~200 fast GPU mins · Adobe: included in Creative Cloud
📅
Subscription (Tiered)
Monthly/annual plans with speed tiers (fast vs relax mode), resolution limits, and commercial rights unlocks. Standard SaaS model.
Runway: $12–$144/mo · Pika: $8–$70/mo · Midjourney: $10–$120/mo
🔌
API Pay-per-use
Charged per image/second-of-video generated. Enterprise and developer-facing. Revenue scales with customer usage.
OpenAI DALL-E 3: $0.04/image (1024×1024) · Stability API: $0.003/step · Runway API: $0.05/sec
🏢
Enterprise Licensing
Custom contracts for volume, commercial indemnification, on-prem deployment, and custom fine-tuning. High ACV ($100K–$5M/yr).
Adobe Firefly Enterprise · Synthesia Enterprise ($24K+/yr) · Runway for Studios
🧩
Platform / Marketplace
Host a marketplace of fine-tuned models (LoRAs, checkpoints). Take a cut of revenue or charge for model hosting. Civitai's model.
Civitai ($3M+ ARR from creator subscriptions) · Hugging Face Hub pro plans
🎯
Vertical SaaS
Embed generation into domain-specific workflows. Higher margins, stickier customers, less competition from horizontal players.
HeyGen (video avatars for sales) · Captions (social creators) · Photoroom (e-commerce)

Unit economics reality check: A100 GPU inference costs ~$2–4/hr. A 1024×1024 image at 20 DDPM steps takes ~0.3 seconds → ~$0.003 compute cost per image. DALL-E 3 charges $0.04 → ~13× markup. Video is 100× more expensive: 5-second 720p clip ~$0.30 compute → Runway charges $0.50–1.00/clip. As model distillation improves (fewer steps), inference costs drop — but so do prices. The sustainable moat is distribution, fine-tuning capability, and commercial rights, not raw generation.

🔬 Active Research Directions

⏱️
Real-Time Generation
Consistency models, SDXL-Turbo, LCM-LoRA reduce steps from 50 to 4. Target: <200ms on mobile. Research in 2025 focuses on flow matching distillation and progressive latent decoding.
🎞️
Long Video Coherence
Maintaining character identity and scene consistency beyond 60 seconds is unsolved. Key papers: FIFO-Diffusion (infinite video), StreamingDiffusion, ring attention for long temporal context.
🕹️
Controllability
ControlNet (pose/depth/edge conditioning) was 2023. Now: IP-Adapter (image prompting), Instant3D (3D consistency), Subject-driven generation (DreamBooth3). 2026 focus: "world models" with physics constraints.
🌐
3D & 4D Generation
Text-to-3D (DreamFusion, Magic3D, Instant3D, Zero-1-to-3) and text-to-4D (video + 3D). NeRF and 3D Gaussian Splatting integration for consistent multi-view generation. Huge applications in gaming, AR/VR, robotics.
🎨
Style & Brand Consistency
IP-Adapter, StyleAligned, Consistent Subject Generation. Enabling brands to lock visual identity across all generated assets. Recraft's "Style" feature is early productization. Key for enterprise adoption.
🔊
Native Audio + Video
Meta's Movie Gen generates video + synchronized audio from text. ElevenLabs + Runway pipelines approximate this. True joint audio-video generation requires temporal alignment across modalities — active research frontier.
🛡️
Watermarking & Provenance
C2PA (Content Credentials) standard for embedding invisible watermarks. Google SynthID, Adobe Content Authenticity Initiative. Critical for combating deepfakes. ICLR 2025 had 20+ papers on invisible watermarking for diffusion models.
🤖
World Models
Video generation as physics simulation. Genie 2 (Google), GameNGen (game simulation), UniSim. The hypothesis: a model that generates physically consistent video has learned a world model — useful for robotics, autonomous driving simulation.

🌐 Applications

🎨
Creative Tools
Concept art, illustration, brand assets. Midjourney, Firefly, FLUX adopted by studios, agencies, solo creators.
📢
Advertising
Product photography, ad creative variants, localisation. Brands use AI to generate 100s of ad variants for A/B testing.
🛍️
E-commerce
Background replacement (Photoroom), model photography (The Fabricant), virtual try-on (Zalando, Amazon). $2B market by 2026.
🎬
Film & Media
VFX assistance, concept visualization, storyboarding, B-roll generation. Runway used in 3 Oscar-nominated productions (2025).
🎮
Gaming
Texture generation, concept art, NPC faces, procedural environment creation. Unity and Unreal both integrating AI generation pipelines.
🏥
Medical Imaging
Synthetic training data generation for rare conditions. Augmentation for CT/MRI datasets. Privacy-preserving patient data synthesis.
🏠
Real Estate
Virtual staging, renovation visualization, exterior rendering. Zillow, Redfin testing AI-generated listing photography.
📚
Education
Interactive diagrams, personalised illustrations, educational video generation. Khan Academy, Duolingo testing AI visual content generation.
👗
Fashion
Virtual fitting rooms, trend visualization, designer concept rapid prototyping. The Fabricant digital fashion house built entirely on generative AI.
🚗
Autonomous Driving
Synthetic training data for rare edge cases (snow, fog, unusual obstacles). Waymo, Tesla, Wayve using video generation for simulation.
📰
News & Publishing
Illustration generation for articles, editorial photos. Associated Press and Reuters using AI image generation under strict editorial guidelines.
🤳
Social & Creators
TikTok, Instagram Reels AI effects. Creator tools for faceless channels. HeyGen, Captions, CapCut dominate the creator economy segment.

Where Does This All Lead?

The generative image and video AI space is moving toward three convergences: (1) modality unification — models that natively understand and generate images, video, audio, and 3D in a single architecture; (2) real-time generation — consistent sub-second image generation and near-real-time video; and (3) world models — where video generation and physics simulation merge into a single learned representation of how the world works.

For builders and investors, the insight is this: the bottleneck is no longer architecture (DiT is good enough). It's data quality, fine-tuning infrastructure, and distribution. The companies that win will be those that build the best data flywheels — products that collect user preference signals, use those signals to improve models, and make the improved models more valuable to users. Midjourney does this with its rating system. Adobe does it with Firefly usage in Photoshop. The open-source ecosystem does it through community LoRA sharing.

We are still in the early phase of a multi-decade transformation in how humanity creates visual content. The tooling you see today — impressive as it is — will look like a sketch compared to what ships in 2028.