Home › Blog › Founding an AGI Startup, Part 7
🔬

The Research Function — Building an R&D Culture on a Startup Budget

DeepSeek shipped a frontier-competitive reasoning model with roughly 160 people — a fraction of an incumbent lab's headcount. What that team's culture actually looked like, and what "research function" should mean for your specific lane, from a full R&D organization to no dedicated research function at all.

FL
FrontierAGI Team
Startups R&D Culture Simulation
What this series is. Part 7 of a standalone founder-playbook simulation, run tri-lane throughout. Team-culture figures below reflect real, sourced reporting on DeepSeek's organization. Educational simulation content, not organizational-design advice for your specific team.

1. 🧭 Does Research Output Actually Scale With Headcount?

The implicit assumption behind most well-funded Lane 1 hiring plans is that more researchers produce proportionally more research output. DeepSeek's own trajectory — a frontier-competitive reasoning model built by a team roughly one-twelfth the size of an incumbent lab's headcount — is the sharpest available counter-example in this entire series, and worth taking seriously as evidence rather than dismissing as an outlier. This article is about what a research function should actually look like at your specific lane and stage, not about maximizing headcount.

~160
DeepSeek's approximate employee count at the time of its R1 release — a fraction of incumbent labs' headcount
~2,000
OpenAI's approximate headcount around the same period — over 12x larger for comparable frontier output
1-2 yrs
Typical experience level of DeepSeek's core technical staff — mostly fresh graduates, deliberately avoiding senior over-hiring
356
Total author count across DeepSeek's seven foundational papers — real, prolific research output from a genuinely small team

2. 🔬 What "Research" Actually Means, By Lane

🔴 Lane 1: Frontier Scale-First

Research is the company — nearly every early hire is a research or research-infrastructure role, and the research function's output (model architecture decisions, training methodology, evaluation frameworks) is the entire product roadmap. The DeepSeek case study below is the central reference point for this lane: it demonstrates real frontier output is achievable well below the headcount incumbent labs operate at, if the culture and hiring philosophy are deliberately built around that constraint rather than treating it as a limitation to overcome by hiring faster.

🔵 Lane 2: Applied / Agentic Layer

Most Lane 2 companies have no dedicated "research" function at all in the traditional sense — and that's correct, not a gap to fill. What functions as research here is closer to rigorous product experimentation: structured A/B testing of prompts and RAG configurations, systematic evaluation of the escalation ladder from Part 6, and staying current on frontier model releases that might change what's technically feasible for your product. A small, dedicated "evals and experimentation" function (even one part-time person) pays for itself quickly, covered in Section 4.

⚪ Lane 3: Narrow Research Bet

Research isn't a function within the company — it's the company's entire reason to exist, and the research culture question is close to existential: does the team have the intellectual honesty to recognize when a specific direction within the broader thesis isn't working, versus the conviction to keep pursuing genuinely hard, multi-year problems through inevitable setbacks? Getting this balance wrong in either direction (too quick to pivot, or too stubborn to update) is a common failure mode specific to this lane.

3. 🌱 Building the Culture — Learning Velocity Over Experience

DeepSeek's specific hiring philosophy — favoring researchers with one to two years of experience over veterans with eight-plus years, with only 20-25% of the team carrying three to five years of experience — is a genuinely different bet than the "hire the most experienced people you can afford" default most well-funded companies pursue. The underlying thesis: at a certain team size, curiosity and learning velocity produce more original research output than experience-heavy hiring at a much larger scale. This isn't a universal prescription — a Lane 1 company without DeepSeek's specific technical leadership and mentorship structure could easily under-hire senior judgment it actually needs — but it's a real, demonstrated alternative to the assumption that more senior hires always produce better research faster.

DeepSeek's bet: curiosity and learning velocity at 150 people can out-produce experience-heavy hiring at ten times the headcount — and the seven-paper, 356-author research output backs that bet up.

4. 📏 Evals as the Research Function Most Lane 2 Companies Actually Need

📊 Structured Prompt/RAG Experimentation
Treating your escalation-ladder decisions (Part 6) as ongoing experiments with measured outcomes, not one-time choices — the closest thing to a "research function" most Lane 2 companies need, and genuinely valuable when done rigorously.
🔍 Independent Benchmark Skepticism
Given how much this entire site's research has shown vendor-reported benchmark numbers can vary by configuration and harness, building your own lightweight evaluation suite specific to your actual use case is more valuable than trusting a frontier lab's marketed benchmark scores.
🔄 Staying Current Without Chasing Every Release
A disciplined, scheduled process for evaluating whether a new frontier model release changes your technical approach — rather than either ignoring new releases entirely or reflexively re-architecting the product every time a lab ships something new.

5. 🏛️ Case Study: DeepSeek's Lean Research Organization

DeepSeek
~160 People, Frontier Output
~150-200 employees at R1 launchMostly 1-2 years experience356 authors, 7 foundational papers
At the time of its January 2025 R1 release — the reasoning model that, as covered in this site's lab-lineage research, triggered real market movement — DeepSeek operated with approximately 150-200 employees, expanding to fill roughly 50 open roles across Hangzhou and Beijing shortly after. This compares to OpenAI's roughly 2,000-person headcount in the same period — over 12x larger for output that was, by the market's own reaction, considered directly competitive. DeepSeek's development model has been described as closer to a traditional academic lab than a typical well-funded startup: learning by working directly on real problems alongside people actively solving them, with most core technical staff carrying just one to two years of experience and the company deliberately avoiding hires with eight-plus years of tenure. Despite the small size, the team's research output was genuinely prolific — 223 researchers were credited across the first five papers, growing to 356 authors across all seven foundational papers.
The lesson for this article: "we need to out-hire the incumbents" is not the only viable research strategy, and for a capital-constrained new entrant, it may not even be the best one — a deliberately lean, learning-velocity-focused culture can produce genuinely frontier-competitive research output, provided the technical leadership and mentorship structure exist to make junior-heavy hiring actually work rather than just being cost-driven under-hiring in disguise.

6. 📋 Side-by-Side: The Research Function by Lane

Factor🔴 Lane 1: Scale-First🔵 Lane 2: Applied Layer⚪ Lane 3: Research Bet
Is there a dedicated research function?Yes — it's the entire companyUsually not, beyond lightweight evalsYes — it's the entire company's purpose
Hiring philosophy optionDeepSeek-style lean/junior-heavy, or traditional senior-heavy — both viable, deliberately chosenNot applicable — no dedicated research hiringTypically senior-heavy, thesis-conviction-driven
What "research output" looks likePublished papers, model releases, benchmark resultsInternal evals, experiment logs, product decisionsPublished papers, eventually a research-validated product (or a validated dead end)
Biggest culture riskOver-hiring seniority without matching learning-velocity cultureTreating "we use AI" as research when it's really just integrationConviction curdling into inflexibility about a failing direction
DeepSeek lesson applicabilityDirectly relevant — a real alternative hiring modelIndirectly relevant — the "lean team, real output" principleDirectly relevant — small, senior research cores follow a similar logic

7. ⚠️ Risk Flags

📈
Assuming Headcount Equals Output
A Lane 1 team that hires aggressively without a matching mentorship and culture structure risks diluting research velocity rather than accelerating it — DeepSeek's case shows the opposite bet can genuinely work, but only with deliberate structure behind it.
🎭
Calling Integration "Research"
A Lane 2 company that labels routine API integration work as "AI research" in its pitch deck or hiring, without the actual experimentation rigor described in Section 4, creates a credibility gap with technically sophisticated investors or hires.
🔒
Thesis Conviction Without an Update Mechanism
Specific to Lane 3 — a research team without an explicit, agreed-upon process for recognizing when a specific sub-direction isn't working can pursue a failing approach far longer than the evidence justifies, burning both capital and investor patience.
🕳️
No Evals at All
The opposite failure mode from over-claiming research — a Lane 2 company with zero structured evaluation of its own product's actual performance is flying blind on the one thing (Part 6's escalation ladder decisions) that most directly affects unit economics and product quality.

8. 🧪 R&D Checklist (All Three Lanes)

1
Decide explicitly whether you need a dedicated research function at all — for most Lane 2 companies, the honest answer is no, and that's the correct answer, not an admission of insufficient ambition.
2
If you're in Lane 1 or 3, choose your hiring philosophy deliberately — DeepSeek-style lean/junior-heavy or traditional senior-heavy — rather than defaulting to "hire the most experienced people the budget allows" without considering the alternative.
3
Build at least a lightweight evals function (Section 4) regardless of lane — even a Lane 2 company benefits from structured, ongoing measurement of its own product's actual performance.
4
Set an explicit thesis-update mechanism for Lane 3 — agree in advance on what evidence would justify changing direction, before the team is emotionally invested in a specific approach that isn't working.
5
Match mentorship structure to hiring philosophy — a junior-heavy team without strong technical leadership isn't replicating DeepSeek's model, it's just under-resourced.

9. 🧭 What's Next in the Series

Part 8 covers Go-to-Market & Marketing — positioning in a crowded field, again across all three lanes: how you get noticed varies enormously between a Lane 2 product competing in the most crowded lane in this entire series, and a Lane 1 or Lane 3 company whose "marketing" is closer to research credibility signaling than a traditional campaign.