Home
›
Blog
›
Building an AGI Startup — Team Profiles & Resumes
Building an AGI Startup · Team Deep Dive
🧑💼
Building an AGI Startup: Who You'd Actually Hire
Sample team profiles and detailed composite resumes for every key role across all three lanes — Applied & Agentic, Frontier Scale-First, and Narrow Research Bet — each paired with a real, named, sourced example of someone actually doing that job in 2026.
FL
FrontierAGI Team
September 10, 2026 · 45 min read
Why Job Titles Don't Tell You Who to Hire
"Founding Engineer" means something different at a Cursor-style applied AI company than it does at a frontier lab training its own model, and something different again at a five-person research bet with no product. Same for "eval owner," "infra engineer," and even "founder." The three-lane framework used throughout this series (Applied & Agentic, Frontier Scale-First, Narrow Research Bet) also predicts almost everything about the right resume for each role — the prior experience that signals fit in one lane can be a mismatch, or even a red flag, in another.
This article makes that concrete. For each lane's core roster, you get: a composite, illustrative candidate profile (a realistic resume built from patterns across many real hires, not any one specific person) and a real, named, sourced example of someone actually holding that kind of role at a real 2026 company. Use the composites as hiring-bar calibration; use the real examples as evidence the pattern is real.
A note on the composite resumes: every "sample candidate" profile below (name, exact background, company history) is an illustrative composite built to be realistic — it does not describe any specific real individual. The real-example boxes are separate, clearly labeled, and sourced to public reporting.
Lane 2 — Applied & Agentic Layer
The Team Building the Product Layer
Lane 2 hiring favors people who have shipped consumer or enterprise software before, who are comfortable with fast iteration against real usage data, and who don't need a research pedigree to be credible. Speed-to-signal and taste dominate over theoretical depth.
Founder / PMLane 2
Composite Profile: "Priya N."
Ex-PM at a Series C fintech app, side-project-to-startup founder
Background: 4 years as a product manager shipping consumer-facing features at a mid-stage fintech startup; built and shipped 2 side projects using LLM APIs that got organic traction on Twitter/Product Hunt before quitting to found full-time.
Technical depth: Can write production TypeScript/Python, is not a systems engineer, relies on a technical co-founder or early hire for architecture decisions.
Product senseRapid prototypingUser interviewsLLM API integrationGrowth loops
Why this fits Lane 2: No research background needed — the job is finding a workflow painful enough that an agent solving it end-to-end is worth paying for, then iterating fast against real usage.
Real Example
Michael Truell — CEO & Co-Founder, Anysphere (Cursor)
One of four MIT classmates who founded Anysphere in 2022; grew Cursor to roughly $1B+ ARR and a reported $60B+ valuation by 2026 by staying personally close to the product — Anysphere is known for having engineers (not dedicated PMs) ship code, talk to users, and drive product decisions directly.
Source: Wikipedia — Anysphere,
Contrary Research.
Founding EngineerLane 2
Composite Profile: "Daniel K."
Ex-Meta infra engineer turned full-stack generalist
Background: 3 years on a backend infra team at a Big Tech company, left to freelance building LLM-powered tools; comfortable across the whole stack (frontend, backend, DevOps) because early-stage agentic products require moving between all three daily.
Technical depth: Deep experience with agent frameworks (LangGraph, custom orchestration), has personally debugged production LLM hallucination/latency issues at scale.
Full-stackLangGraph/CrewAISandboxed execution (E2B)Prompt engineeringProduction debugging
Why this fits Lane 2: Generalist range matters more than specialization — the first 5-10 engineers at an applied AI startup each own an entire vertical slice of the product, not one narrow function.
Real Example
Sualeh Asif & Aman Sanger — Co-Founders, Anysphere
MIT classmates of Michael Truell who became CPO and COO respectively; both were technical co-founders who wrote code directly on Cursor's early product rather than being hired into a purely managerial role — illustrating how Lane 2's "founding engineer" archetype often is the founding team itself in the first year.
Source: Wikipedia — Anysphere.
DesignerLane 2
Composite Profile: "Elena V."
Ex-product designer at a dev-tools startup
Background: 3-4 years designing developer-facing tools (IDEs, CLIs, dashboards); strong opinions on how to make agent autonomy legible to a human user — when to show reasoning traces, when to ask for confirmation, when to just act.
Technical depth: Can prototype in code (React), works directly in the codebase rather than handing off static mockups — near-universal at this team size.
Interaction designDeveloper UXDesign systemsFrontend prototypingTrust & transparency UI
Why this fits Lane 2: The central design problem in agentic products — how much autonomy to grant an agent by default and how to surface its actions — barely exists in Lanes 1 or 3, making this a genuinely lane-specific skill.
Real Example
Cursor's Product Design Practice
Reporting on Anysphere's ~300-person team culture describes engineers and designers working without a layer of dedicated product managers, with design decisions (like Cursor's Tab autocomplete and agent-mode UI) made in tight loops between design and engineering rather than through a formal handoff process.
Source: JobsByCulture — Working at Cursor 2026.
Eval / QA OwnerLane 2
Composite Profile: "Marcus T."
Ex-ML engineer specializing in evaluation tooling
Background: 2-3 years building internal eval harnesses at an ML-heavy company; understands both classical software QA and LLM-specific failure modes (hallucination, tool-call errors, task drift in long agent loops).
Technical depth: Fluent with eval frameworks (DeepEval, Anthropic's Inspect AI), builds custom eval suites tied directly to the product's specific failure modes rather than generic benchmarks.
Eval harness designDeepEval / Inspect AIAgent trace analysisRegression testingLangSmith / Arize
Why this fits Lane 2: Because agentic products fail in open-ended, hard-to-predict ways once deployed, a dedicated eval owner is now considered close to a first-10-hires role rather than a later addition — a shift from just 2 years ago.
Real Example
Anthropic's Inspect AI & Adopter Ecosystem
Anthropic's own open-source Inspect AI framework is explicitly built for the eval-owner function this profile describes, and is now widely adopted by applied AI startups (per this series' Lane 2 article) as the default tool for exactly this role rather than a bespoke in-house system.
Source: Inspect AI (UK AISI / Anthropic ecosystem).
First Sales / GTM HireLane 2
Composite Profile: "Jordan P."
Ex-enterprise AE at a vertical SaaS company
Background: 3-5 years closing mid-market/enterprise deals at a SaaS company selling into a similar buyer persona; comfortable running technical proof-of-concept pilots directly with engineering teams, not just procurement.
Technical depth: Not an engineer, but fluent enough in the product's underlying agent architecture to answer technical objections in a pilot conversation without escalating every question to engineering.
Enterprise salesTechnical POCsPilot-to-contract conversionUsage-based pricing
Why this fits Lane 2: Applied AI companies sell primarily on demonstrated task success in a pilot, not on a slide deck — the first sales hire is closer to a solutions engineer than a classic AE.
Real Example
Decagon's Founder-Led Enterprise Motion
Decagon, founded by Jesse Zhang and Ashwin Sreenivas (ex-Scale AI), reached a $4.5B valuation with real enterprise revenue while its technical founders stayed personally involved in technical reviews during the sales process — the founder-as-early-salesperson pattern common across Lane 2 before a dedicated sales org is built out.
Source: Sacra — Decagon.
Infra / Reliability EngineerLane 2
Composite Profile: "Wei C."
Ex-SRE at a high-traffic API company
Background: 4 years running production reliability for a high-QPS API product; comfortable with sandboxed code execution environments and the specific reliability challenges of agent loops that can run for minutes rather than milliseconds.
Technical depth: Hands-on with sandboxing/execution infra (E2B, Modal, Browserbase), understands cost-per-agent-run economics well enough to catch runaway loops before they hit the bill.
E2B / Modal / BrowserbaseSandboxed executionCost/latency tradeoffsObservability
Why this fits Lane 2: Reliability here means bounding an agent's real-world side effects (file writes, API calls, purchases) — a different discipline from classic web SRE work, even though the resume pattern looks similar on paper.
Real Example
E2B, Browserbase & Modal Labs' Customer Base
These three infrastructure providers, covered in this series' Lane 2 article, exist specifically because applied AI startups need dedicated sandboxed-execution reliability rather than building it from scratch — most Lane 2 infra hires now spend their first months integrating one of these rather than building bespoke sandboxing.
Source: E2B.
Lane 1 — Frontier Scale-First
The Team Training a Foundation Model
Lane 1 hiring is dominated by prior frontier-lab experience — DeepMind, OpenAI, Meta FAIR, Anthropic — because the specific challenges of large-scale distributed training and pre-training research are only learned by doing them at that scale before. Resumes here read very differently from Lane 2's.
Founder / Research LeadLane 1
Composite Profile: "Dr. Anaïs R."
Ex-DeepMind research scientist, PhD in ML from a top European lab
Background: 6 years at a frontier lab contributing to large-scale pretraining projects; left with 2-3 co-founders from the same lab to pursue a specific, differentiated training thesis (e.g. sovereignty, open-weight distribution, or data-efficiency).
Technical depth: Deep, hands-on pretraining and scaling-law experience; the resume itself is the primary fundraising asset in this lane.
Large-scale pretrainingScaling lawsDistributed trainingFundraising credibility
Why this fits Lane 1: Capital access in this lane is underwritten almost entirely by founder research pedigree — a founder without this background struggles to raise the $200M+ needed regardless of business plan quality.
Real Example
Arthur Mensch — CEO & Co-Founder, Mistral AI
Spent years at Google DeepMind contributing to major pretraining projects including Retro and Chinchilla before co-founding Mistral in 2023 with fellow ex-Meta FAIR researchers Timothée Lacroix and Guillaume Lample — Mensch's DeepMind pretraining pedigree was central to Mistral raising toward an $830M+ round and becoming Europe's leading sovereign AI challenger.
Source: AI Wiki — Arthur Mensch.
Distributed Systems / Infra EngineerLane 1
Composite Profile: "Kenji O."
Ex-infra engineer on a large training cluster team
Background: 4-5 years working on multi-thousand-GPU training infrastructure at a frontier lab or cloud provider; has personally debugged interconnect failures and checkpoint-recovery issues during live, multi-week training runs.
Technical depth: Fluent in Megatron-DeepSpeed/FSDP, understands neocloud compute economics (spot vs. reserved, interconnect topology) well enough to negotiate a compute deal, not just operate one.
Megatron-DeepSpeedFSDPMulti-node GPU clustersCheckpoint recovery
Why this fits Lane 1: This is the single scarcest skill in the lane — a research-heavy, infra-light team bottlenecks exactly here, as flagged as a risk in this series' Lane 1 article.
Real Example
Reflection AI's Distributed Systems Engineer Role
Reflection AI's own public job listing for "Member of Technical Staff — Distributed Systems Engineer" describes exactly this profile, underscoring how central this hire is to a lab that just closed a $1B Nebius compute deal and needs infra talent to actually use it.
Source: Reflection AI job posting (Ashby).
ML Research ScientistLane 1
Composite Profile: "Sofia B."
PhD in ML, postdoc or research-engineer stint at a frontier lab
Background: PhD + 2-3 years as a research engineer on pretraining or architecture research at a major lab; publication record in top ML venues signals both technical depth and the ability to attract further research talent.
Technical depth: Comfortable owning a full research thread (data mix ablations, architecture variants) semi-independently within the pretraining pipeline.
PyTorch/JAXData mix ablationsArchitecture researchDistributed training (Ray)
Why this fits Lane 1: Mistral's own hiring bar for this role — a Master's/PhD plus 4+ years on large-scale ML codebases — reflects the field-wide norm that this is not an entry-level position at any credible frontier lab.
Real Example
Pavankumar Reddy M. — Research Scientist, Mistral AI
A publicly listed research scientist at Mistral AI, illustrative of the profile Mistral's own published job requisitions describe: large-scale ML codebase experience, PyTorch/JAX fluency, and distributed-systems familiarity (Ray/Kubernetes) as baseline requirements even for individual-contributor research roles.
Source: LinkedIn (public profile),
Mistral job posting.
Post-Training / Alignment ResearcherLane 1
Composite Profile: "Tomás F."
Ex-RLHF researcher, background spanning both capabilities and safety teams
Background: 3-4 years working on RLHF/preference-tuning pipelines at a lab that has shipped at least one publicly released model; has hands-on experience with the specific failure modes that show up only after a model is post-trained and evaluated broadly.
Technical depth: Fluent across SFT, RLHF/DPO, and red-teaming methodology; often the bridge between the research core and the eval/safety function.
RLHF / DPOSFT pipelinesRed-teamingPreference data curation
Why this fits Lane 1: Post-training is where a raw pretrained model becomes a usable, safe product — a lab that underinvests here ships a technically impressive but commercially unusable model.
Real Example
John Schulman — Chief Scientist, Thinking Machines Lab
Co-founder of OpenAI who helped architect ChatGPT's RLHF post-training pipeline, later joined Thinking Machines Lab as chief scientist in Feb 2025 — a direct real-world example of post-training research pedigree being the deciding hire for a Lane 1-adjacent lab's credibility.
Source: Wikipedia — John Schulman.
Safety / Eval Team LeadLane 1
Composite Profile: "Ingrid L."
Ex-safety researcher at a frontier lab, background in applied AI research
Background: 3+ years building or contributing to a safety/eval team at a major lab from an early stage; has personally built internal eval infrastructure rather than only using existing benchmarks.
Technical depth: Fluent with HELM/Inspect AI-style eval frameworks, comfortable owning both capability evals and safety red-team processes simultaneously at a small lab's scale.
HELM / Inspect AIRed-teamingSafety eval infraModel card authorship
Why this fits Lane 1: A frontier lab claiming credibility on safety needs someone who has actually built that function before, not just studied it academically.
Real Example
Lilian Weng — VP of Research (Safety), Thinking Machines Lab
Spent years at OpenAI building its applied AI research and safety functions from scratch before co-founding Thinking Machines Lab in Feb 2025 as VP of Research focused on safety — a direct example of a safety-team-lead hire whose entire prior career was building that exact function at a frontier lab.
Source: American Bazaar.
Lane 3 — Narrow Research Bet
The Team Betting on One Research Thesis
Lane 3 hiring is the smallest and most senior of the three lanes by design — teams stay deliberately lean, and every hire is expected to carry independent research judgment rather than execute a defined roadmap.
Founder / Research LeadLane 3
Composite Profile: "Dr. Yusuf A."
Former Chief Scientist-level researcher at a frontier lab
Background: 8-10+ years at the very top of frontier AI research, with a body of work the entire field recognizes; leaves to pursue one specific, narrow thesis they believe the industry has gotten wrong.
Technical depth: The resume itself, not a business plan or product demo, is what raises the round — this lane is underwritten almost entirely on this single factor.
Field-defining research recordThesis articulationTalent magnetism
Why this fits Lane 3: Per this series' idea-scoring framework, founder/team research credibility is close to a precondition for raising in this lane at all.
Real Example
Ilya Sutskever — CEO & Co-Founder, Safe Superintelligence Inc.
Former Chief Scientist and co-founder of OpenAI, whose personal research credibility is widely cited as the single reason SSI could raise $8B at a $32B valuation with zero shipped products — the purest real-world case of founder pedigree substituting entirely for product or revenue signal.
Source: Wikipedia — Safe Superintelligence Inc..
Senior Research ScientistLane 3
Composite Profile: "Dr. Hana S."
PhD + 5-7 years at a top lab, joined pre-thesis-validation
Background: Strong independent publication record; joined the research bet specifically because they share the founder's belief the thesis is underrated, not for compensation or brand — equity and mission alignment matter more here than at any other lane.
Technical depth: Capable of designing and running an independent experimental thread with minimal oversight; often has a prior working relationship with the founder.
Independent research designFast experiment iterationThesis-driven judgment
Why this fits Lane 3: A ~50-person team at a $32B valuation (SSI's real ratio) only works if each hire operates with near-total independent research judgment rather than needing management overhead.
Real Example
Daniel Levy — President & Co-Founder, SSI
PhD in computer science from Stanford, prior senior AI engineer role at OpenAI with a strong safety focus, plus earlier technical internships at Microsoft, Meta, and Google — became SSI's president after co-founder Daniel Gross departed for Meta in mid-2025, illustrating the depth of individual research credibility this lane requires even in its second-most-senior seat.
Source: HyScaler.
Interpretability / Domain Research ScientistLane 3
Composite Profile: "Dr. Owen M."
Ex-DeepMind interpretability researcher, founding a specialized research team
Background: Founded or co-founded an interpretability research function at a major lab before joining or founding a narrower, more focused research bet applying those methods outward (e.g. to biology or materials).
Technical depth: Deep mechanistic-interpretability method expertise, plus willingness to partner directly with domain scientists outside ML entirely.
Mechanistic interpretabilityModel probing/steeringCross-domain collaboration
Why this fits Lane 3: Domain-science neolabs specifically need this rare combination — deep ML interpretability skill plus genuine comfort partnering with non-ML domain experts.
Real Example
Tom McGrath — Chief Scientist & Co-Founder, Goodfire
Founded Google DeepMind's interpretability team while completing his PhD, then co-founded Goodfire in June 2024 as Chief Scientist — Goodfire went on to identify a novel class of Alzheimer's biomarkers by applying interpretability techniques to a biology foundation model built by domain partner Prima Mente, the exact cross-domain pattern this profile describes.
Source: Lightspeed Venture Partners — Goodfire.
Founding Research EngineerLane 3
Composite Profile: "Nadia K."
Ex-research engineer, builds the tooling researchers need rather than the researcher's own experiments
Background: 2-4 years as a research engineer supporting a scientist-heavy team; builds and maintains the internal experiment infrastructure, interpretability tooling, or eval harness the research core needs but doesn't want to build themselves.
Technical depth: Strong software engineering combined with enough ML fluency to understand what the research team actually needs, often the difference between a productized platform (like Goodfire's Silico) and one that never leaves the lab.
Research infra toolingInternal platform buildingML fluency
Why this fits Lane 3: This is often the first hire that isn't purely a senior researcher — a signal the thesis has moved from Stage 1/2 (formation/validation) toward Stage 3+ (scale-up) in this series' Lane 3 pipeline.
Real Example
Dan Balsam — CTO & Co-Founder, Goodfire
Previously founding engineer and Head of AI at RippleMatch, where he built the core engineering organization and deployed LLMs in production, before co-founding Goodfire as CTO — the engineering-builder counterpart to McGrath's research leadership, responsible for turning Goodfire's interpretability research into the productized Silico platform.
Source: Contrary Research — Goodfire.
Founder / Business & Partnerships LeadLane 3
Composite Profile: "Marcus B."
Ex-applied AI startup founder, pivots into a supporting role at a research-heavy company
Background: Previously founded or led go-to-market at an applied AI company; joins a research-bet lab specifically to own fundraising logistics, compute partnerships, and (eventually) the first productization effort — deliberately not trying to direct the research itself.
Technical depth: Enough ML fluency to translate research progress into a fundraising or partnership narrative, without needing to be a researcher themselves.
Fundraising narrativeCompute partnershipsProductization strategy
Why this fits Lane 3: Even the purest research bets eventually need someone translating research momentum into the capital and compute needed to keep validating the thesis — usually a business-minded co-founder rather than a hired executive.
Real Example
Eric Ho — CEO & Co-Founder, Goodfire
Previously co-founded RippleMatch, an applied AI startup, before co-founding Goodfire in June 2024 as CEO — bringing product, engineering, and go-to-market experience to complement McGrath's pure research pedigree, a direct real-world case of the business/partnerships-lead profile pairing with a research-first co-founder.
Source: Longterm Wiki — Goodfire.
Cross-Lane Comparison: Same Title, Different Resume
| Archetype | Lane 2 Resume Signal | Lane 1 Resume Signal | Lane 3 Resume Signal |
| Founder | Prior product/PM experience, shipped consumer or B2B software fast | Frontier-lab pretraining research credibility, capital-raising track record | Field-defining individual research record (Chief-Scientist-level or above) |
| Core builder | Full-stack generalist, agent-framework fluency, fast iteration instinct | Distributed-systems/infra depth at multi-thousand-GPU scale | Independent research judgment, minimal need for oversight |
| Support function | Designer fluent in agent-autonomy UX; sales hire who runs technical pilots | Post-training/RLHF researcher; dedicated safety-eval team lead | Founding research engineer building internal tooling; business/partnerships co-founder |
| Team size at scale | 50-300+ staff is normal (Cursor: ~150-300) | Hundreds to low thousands (Mistral: ~450) | Deliberately lean — SSI: ~50 staff at a $32B valuation |
| What raises the next round | Usage growth, revenue, retention metrics | Model benchmark parity, compute partnerships, research credibility | Founder/team pedigree alone, often with no product at all |
The same word — "founding engineer" — describes a full-stack generalist shipping agent UX in Lane 2, a distributed-systems specialist keeping a training run alive in Lane 1, and a research-engineer building internal tooling for five PhDs in Lane 3. Reading a resume without first knowing which lane you're hiring for is close to reading it blind.
How to Read a Resume Differently by Lane
🧭 Quick Diagnostic
- In Lane 2, weight shipped-product velocity and usage-data fluency over academic credentials — a candidate who has never published a paper but has shipped three real products fast is often the stronger hire.
- In Lane 1, weight specific frontier-lab pretraining or distributed-systems experience heavily — general "AI experience" without hands-on large-scale training exposure is a weak signal here specifically.
- In Lane 3, weight independent research judgment and thesis alignment over almost everything else — a stellar resume that doesn't genuinely believe the specific thesis is a worse hire than a slightly less decorated one who does.
- Across all three lanes, the resume that looks best on paper for a different lane is often a subtle red flag — a frontier-lab pretraining researcher joining a Lane 2 applied startup, or a fast-shipping Lane 2 product engineer joining a Lane 1 training team, both require an honest conversation about why the lane switch makes sense.
🔗 Sources for Real Examples Cited Above
🧭 Closing — Hire for the Lane, Not the Title
🎯 The Bottom Line
Across all three lanes of this series, the single biggest hiring mistake is importing a resume pattern that worked in a different lane — a brilliant frontier-lab researcher is not automatically a good Lane 2 founding engineer, and a fast-shipping product engineer is not automatically ready for a multi-week distributed training run. The real, named examples above — from Anysphere's founder-engineers to Mistral's DeepMind-trained research team to SSI's Chief-Scientist-level founding core — all share one trait: their prior experience maps almost exactly onto what their specific lane actually requires, not onto AI experience in general. Whichever lane you're building in, use this as your hiring bar: does this resume signal the specific pattern that lane rewards, or just "AI experience" in the abstract?