Distillation, Centralization, and the Moat Question
If recursive self-improvement lets one lab pull ahead, does that lead stay locked inside the walls that produced it โ or does distillation let everyone else catch up within months? This piece separates what's actually confirmed about the OpenAIโDeepSeek distillation dispute from what's still just allegation, walks through what DeepSeek's $5.6M figure does and doesn't cover, and looks at the best available quantified answer โ from Epoch AI โ to how much of a lead a frontier lab actually keeps.
A Lead Is Only Worth as Much as It's Hard to Copy
Parts 1 and 4 of this series described real, verified systems โ Dream-RSI, ModularRSI, ScienceBuddy, AlphaEvolve โ that let a frontier lab's own AI accelerate its own research. Suppose one of those labs gets meaningfully ahead this way: faster kernels, better agent scaffolds, a research loop that compounds. The natural next question is economic, not technical: does that advantage stay inside the lab that built it, or does it leak out to competitors within months, through the very same mechanism โ training a new model to imitate an old one's outputs โ that made distillation a routine technique in the first place?
This question became concrete, not hypothetical, in January 2025, when DeepSeek released R1 at a small fraction of the reported cost of comparable Western models, and OpenAI publicly suggested DeepSeek had used exactly this technique against OpenAI's own models. What follows is an attempt to separate the parts of that story that are actually confirmed from the parts that remain allegation โ because the two get conflated constantly in casual retellings, and the distinction matters for whether "distillation erases the moat" is a proven pattern or a plausible but unproven one.
What Distillation Actually Is
Knowledge distillation โ training a smaller "student" model to reproduce a larger "teacher" model's outputs โ is a decade-old, entirely ordinary technique, not a shortcut invented to steal frontier labs' work. It was formalized by Geoffrey Hinton, Oriol Vinyals, and Jeff Dean in 2015 as a way to compress a large model's knowledge into a smaller, cheaper one that keeps most of the teacher's performance. Every major lab uses some version of it.
Meta's Llama 3.2 technical documentation explicitly describes the 1B and 3B parameter models as distilled from the larger Llama 3.1 8B and 70B models. That is the clearest, most directly lab-confirmed example available. GPT-4o-mini and Gemini Flash are widely reported and generally assumed by outside analysts to be distilled from their larger siblings (GPT-4o and Gemini Pro respectively) โ but neither OpenAI nor Google has published a technical document explicitly using the word "distilled" to describe them, so that specific claim should be read as informed inference, not lab-confirmed fact.
None of this is controversial when a lab distills its own model into a smaller version of itself. The DeepSeek dispute is about something different: whether an outside party can distill a competitor's closed model by training on large volumes of its API outputs โ effectively using the competitor's own model as an unpaid teacher.
The DeepSeek Accusation, Carefully Separated From What's Proven
Here is the timeline as it's actually documented, with the allegation clearly marked as an allegation:
It's worth sitting with why this matters for the moat question specifically. If the accusation were proven, it would show that even a closed, API-gated frontier model isn't fully protected from distillation by outsiders โ anyone with enough API budget and cleverness about data collection could, in principle, extract a meaningful fraction of a teacher model's capability. That would be a strong data point for "moats are weak." Because it isn't proven, the strongest honest claim is narrower: the accusation was plausible enough that OpenAI staked its reputation on making it, which itself says something about how seriously frontier labs take this risk โ but it isn't evidence the risk materialized.
What DeepSeek's $5.6M Figure Actually Covers
The other half of the "moats are collapsing" narrative rests on DeepSeek-V3's widely repeated training cost: $5.576 million. That figure is real โ it appears in DeepSeek's own technical report, calculated from 2.788 million H800 GPU-hours at an assumed $2-per-GPU-hour rental rate. But the report itself is explicit about scope, and that scope is almost always dropped when the number gets repeated:
DeepSeek's own report says the $5.576M figure explicitly excludes costs from prior research and ablation experiments on architecture, algorithms, and data, and excludes the underlying hardware capital expenditure entirely. SemiAnalysis's independent analysis, using DeepSeek's estimated compute fleet (widely reported in the low tens of thousands of Hopper-class GPUs, with exact counts disputed across sources), put total infrastructure investment closer to the billions, not millions.
How Fast Does the Capability Gap Actually Close?
Rather than relying on any single model's story, Epoch AI has tried to measure the lag between frontier closed-weight models and the open-weight models that eventually reach similar capability, using their Epoch Capability Index (ECI). Their October 2025 analysis found open-weight models lagged the closed-weight frontier by an average of about three months, measured across the period from January 2023 to October 2025. A later update found the gap had widened slightly to roughly four months (about 8 ECI points) โ comparable to the capability difference between GPT-5 and GPT-5.5.
This is a meaningfully different picture from either extreme in the public debate. It isn't "moats don't exist" โ a three-to-four-month lead is real and, in a fast-moving field, can matter a great deal for who captures a given wave of enterprise adoption or research attention. It also isn't "moats are durable for years" โ three to four months is a short enough window that a lab planning around a multi-year technical lead is planning around something the evidence doesn't currently support.
Correcting Our Own Record
What Greenblatt has actually published, in a Redwood Research blog post from around June 2025 titled "AI safety techniques leveraging distillation," is narrower and framed differently than a competitive-moat argument. His point is that because it's currently cheap to train a new model to imitate a strong model's outputs, that cheap reproducibility can be turned into a safety tool โ for example, distilling a model's behavior into a fresh copy can help researchers isolate and study specific behaviors, including potentially misaligned ones, more tractably than probing the original model directly. That's a genuinely different argument from "distillation means no lab keeps a durable advantage," and this article should have kept those separate from the start.
The Compute and Capital Moat Debate
Beyond the specific DeepSeek case, there's a broader, ongoing public argument about whether compute and capital create a durable moat at all. This is genuinely contested territory, and โ in the spirit of the correction above โ this article will describe the shape of the debate without over-claiming specific quotes it hasn't independently verified against primary sources.
- One camp argues capability commoditizes quickly: once a capability level is demonstrated to be possible, competitors and open-source projects can typically reach it within a matter of months (consistent with the Epoch AI figures above), meaning the durable moat, if one exists, has to come from something other than raw model capability โ distribution, brand, enterprise integration, proprietary data, or compute access at a scale that's hard to replicate quickly.
- The other camp points to the sheer capital intensity of frontier training runs and the compounding advantage of owning the infrastructure, the talent density, and the iteration speed that lets a lab try more ideas per unit time โ an advantage that a single distilled model release doesn't erase, because the next frontier run is already underway by the time competitors catch up to the last one.
Why This Matters Specifically for Recursive Self-Improvement
If a lab achieves a genuine RSI-driven research-speed advantage โ say, an internal system that measurably accelerates its own experiment cycle the way Dream-RSI, ModularRSI, and ScienceBuddy (covered in Part 4) are already starting to do in narrower ways โ the Epoch AI data above suggests that advantage would likely show up in competitors' models within a few months, not stay locked away for years. That has two very different safety implications depending on which failure mode you're more worried about:
- If You Worry About Concentration of Power
- Fast diffusion is reassuring โ no single lab gets to keep an RSI-driven advantage to itself for long enough to translate it into unchecked unilateral power.
- If You Worry About Coordination and Safety Margins
- Fast diffusion is alarming โ it means any safety-relevant capability jump gets replicated by every other lab (including ones with weaker safety practices) within months, compressing the time available to notice and respond to a problem before it's everywhere at once.
This series takes no position on which risk is larger โ that's precisely the kind of expert-disputed judgment call the next section is about โ but the Epoch AI lag data is the closest thing available to a quantified answer for how much time either camp should assume they actually have.
Expert Disagreement
Three Parts Published, Three to Go
Readiness Checklist
โ ๏ธ What's Missing or Uncertain
๐ Reference Links
- Forbes โ "Did DeepSeek Copy Off of OpenAI, and What Is Distillation?" (Jan 30, 2025)
- South China Morning Post โ coverage of OpenAI's distillation allegation
- Epoch AI โ Open-weight models lag the closed-weight frontier by ~3 months
- Epoch AI โ updated open/closed capability-gap analysis (~4 months, 8 ECI points)
- SemiAnalysis โ "DeepSeek Debates": independent estimate of DeepSeek's true infrastructure cost
- Redwood Research โ Ryan Greenblatt, "AI Safety Techniques Leveraging Distillation"
- This site โ The DeepSeek Lineage: A Consolidated Comparison
- This site โ Recursive Self-Improvement: A Primer (Part 1)
- This site โ AI Automating AI Research: RSI's Testing Ground (Part 4)