Blog

The Neolab Wild West

by Thomas Joshi, Madison Faulkner, Lila Tretikov and Andrew SchoenJun 17, 2026

Exiting the Frontier Labs Era for AI Venture Capital

Venture has traditionally focused on revenue scale and growth when evaluating large capital raises. However, in the past few months, the traditional AI venture model has been upended by multibillion-dollar rounds for companies with no revenue, no product, and in some cases no model. The question is why. The answer starts with the scale of what has already been funded— AI research labs now represent roughly $1.68 trillion of aggregate private valuation.1 

And 93.6% of that $1.68 trillion sits in a single bucket: frontier generalists2, with OpenAI, Anthropic, xAI, and Safe Superintelligence absorbing almost all of it. Every other segment a venture investor might back including biotech, robotics, scientific discovery, voice, edge, and more combines to barely 6%. The Neolab market, in other words, is not a market—it’s one concentrated basket with four names.

Every dollar that flows into a frontier generalist at hundreds of billions in valuation assumes that their next model will deliver another step change in model performance to temporarily capture a market held by a competitor in an endless tug of war. VCs are not underwriting today's models. They are assuming that scale, plus a breakthrough no one has seen, will manifest into market creation and capture.

We wonder where the next leg of AI venture returns are made.

Specialization Defining AI Innovation Trends

There are three key trends that point to a potential direction for the next generation of innovative AI companies:

  1. Commoditization. Open-weight models are now closing the capability gap at a fraction of the compute that closed-source labs spent to open it. Training data per active parameter has grown 3.1x per year since 2022, dataset sizes double roughly every six months, and estimations suggest training runs longer than nine months become structurally inefficient sometime around 2027.3 The marginal pre-training dollar is buying less capability than the dollar before it. Incremental performance gains will keep moving market share at the margins, but only massive innovation on a specific use case or modality actually has any potential to displace an incumbent at this stage. This new innovation could look like a verifiable reasoning loop, a new modality, or a domain-native architecture. The most visible work today sits one layer above the weights at the behavior and alignment layer, where labs are fine-tuning specifically for productivity use cases. For example, Sovereign AI Provider Reflection AI will serve as the AI model provider to the U.S. National Labs. Specifically, they are partnering on the Genesis Mission, which is a Department of Energy initiative to accelerate scientific research through AI.4 What separates the top models is no longer what they know but how they act within specific environments. 

  1. Shift in budget. As we have shifted to an RL-first (Reinforcement Learning-first)  paradigm where a system is trained primarily through environmental rewards and trial-and-error, rather than relying heavily on human-labeled data or pre-existing templates, the question isn’t just who has the largest cluster, but also who has the best environments, curricula, and verifiers? Compute is no longer the only moat; the entire loop is a moat. The foundation labs are not sufficiently focused on the correct product abstraction that makes RL deliver business outcomes and have instead curated RL environments for diverse, disjointed tasks without focused taste.

Grok’s training shows a shift toward RL

  1. The exit. Top-decile software now trade at ~15-18x next-12-months revenue against ~3-6x for the rest of the index, and strategic M&A volumes have climbed back to 2022 highs. The strategics are paying premium multiples for AI-native IP they cannot build internally fast enough, and they are buying specific capability, not general intelligence.

These trends lead us to one conclusion. The next generation of AI companies that deliver venture-like outcomes will likely be Neolabs: research-first companies that own a domain-biased corpus, run an integrated RL and verification loop perhaps inside a high-stakes domain, and could exit into public markets or a strategic M&A wave that is already underway. SAP's May 2026 acquisition of Prior Labs is the clearest recent example: a strategic with its own in-house AI effort committing more than €1B to bring in a specialist team because tabular foundation models are not a capability a general LLM does well enough on.

Why Neolabs' Framework Represent a New Venture Capital Opportunity

The next jump in performance requires a continuous training run longer than nine months, weighed against the opportunity cost of a frozen flagship model. 

Some argue this is temporary: The next architecture will reset the curve and restore the frontier's lead. The history of foundation models suggests otherwise. Every architectural advance of the last three years, from mixture-of-experts to long-context attention to reasoning-trace distillation, has been replicated in open weights.

The durable wedges are domain-specific data that no one else can assemble, workflows embedded so deep in a customer's operations that ripping them out would cost more than keeping them, and distribution channels that the hyperscalers cannot easily replicate. Which is why labs like OpenAI have spent billions starting partnerships such as the OpenAI Deployment Company.5

The Three-Layer Framework: Where AI Neolab Value Gets Created

Neolabs can be built around three, specific layers, which also act as moats.

  1. The first layer is the corpus. 

Recent CMU work on mid-training shows that a model pre-trained on a domain-biased corpus, then fine-tuned with reinforcement learning, outperforms the same model trained with RL alone by margins that grow with task difficulty. If a base model is broad and opinionated about everything, post-training has to fight a vast prior built from Reddit, fiction, and the open web. If the base is already fluent in protein structures, tabular data, or German legal code, every subsequent RL run is cheaper, faster, and more stable. A domain-biased corpus is a one-time investment that pays back on every training cycle after it. This is why Xaira, Isomorphic Labs, and Chai Discovery are not "biotech startups using AI" but biotech-native foundation model labs, why Prior Labs built TabPFN as a foundation model purpose-built for tabular data rather than another general LLM with a spreadsheet wrapper, and why sovereign labs like Aleph Alpha and Sarvam matter. Their corpora cover regulated and linguistic domains that OpenAI cannot legally or practically assemble. 

Research from CMU shows Mid-training+RL beats RL alone

  1. The second layer is the RL loop. 

Whoever owns the environment owns the data flywheel, because every action the model takes inside that environment generates a new training signal that no one outside the environment can replicate. CuspAI builds materials-discovery environments where simulated chemistry is the training signal. Factory AI does the same thing one domain over, running continual learning at the coding interface so that every mission a customer runs captures reusable patterns into a skill library. Vertical specific partnerships do not just require architectural change, they also require an entire organizational change.

  1. The third layer is the verifier.

Reinforcement learning only works when the reward signal is trustworthy, and in most domains it is not. Verifiers like "Did the molecule bind to the target with sub-nanomolar affinity" or "Did the synthesized material exhibit the predicted band gap" include deterministic rewards, and they are only available in domains where ground truth exists and someone has built the apparatus to measure it. The verifier is the most underestimated asset in this stack because it looks like infrastructure when it is actually the source of truth.

The three layers reinforce each other. The corpus makes RL tractable. The RL loop generates the data the corpus cannot. The verifier ensures both are pointed at the correct target. A frontier generalist optimizes one layer at a time across every domain at once and ends up dominant in none. A Neolab integrates all three inside a single domain.

The Neolab opportunity, then, sits exactly where frontier generalists were never built to win: fields with proprietary corpora difficult to assemble, environments they do not natively operate inside, and verifiers only the practitioners can craft.

Why Partner with NEA: Funding New AI Innovation 

We know what it takes to build and deploy frontier AI research given our experience as a co-founder of DSPy and the former Deputy CTO of Microsoft during the formation of the Microsoft–OpenAI partnership. We have also backed a select number of Neolabs and continue to invest in them. We believe these segments will produce a meaningful share of the next trillion dollars of Neolab value.

Another reason is scale. Neolabs raise $50M to $500M before they have revenue, which means their investors have to be comfortable holding the position for five to seven years before a commercial signal arrives. Their partners must then have sufficient reserves to continue to back that company through every stage in their journey. Only a few firms in the world have the long-term orientation, culture, and track record to do that.

And finally, technical expertise. NEA has backed World Labs, Fei-Fei Li's lab building spatial intelligence models for the three-dimensional world; Sakana AI, the Tokyo-based lab applying nature-inspired methods to model architecture; and CuspAI, the materials-discovery lab where simulated chemistry is the training signal. Each is now a leader in its category. The lessons learned from these investments are carried into partnerships with the next round of labs. 

If you are a founder creating a Neolab or AI Researcher/Engineer looking to join one, please reach out to us: tjoshi@nea.com, ltretikov@nea.com, aschoen@nea.com, mfaulkner@nea.com.

AI21 Labs

Description: Enterprise LLM developer building summarization, rewriting, and natural-language comprehension tools.

Category: Frontier Generalist

Anthropic

Description: Safety-focused frontier AI lab building Claude and the alignment research that underpins it.

Category: Frontier Generalist

DeepMind

Description: Alphabet’s frontier AI research lab building Gemini and the underlying research across language, vision, and embodied AI.

Category: Frontier Generalist

Humans&

Description: Building AI systems that interact collaboratively with people.

Category: Frontier Generalist

Showing 4 of 86

About the Authors

Thomas Joshi

Thomas joined NEA in 2025 as an Associate on the Technology Investing Team focused on early and growth-stage investments. Previously, he held several AI Researcher, AI Engineering, and finance positions. Thomas is co-author of Stanford DSPy, the most popular open source software to come out of Stanford AI has been used by companies like Meta and Microsoft. Thomas graduated from Columbia University with a B.Sc. in Artificial Intelligence.
Thomas joined NEA in 2025 as an Associate on the Technology Investing Team focused on early and growth-stage investments. Previously, he held several AI Researcher, AI Engineering, and finance positions. Thomas is co-author of Stanford DSPy, the most popular open source software to come out of Stanford AI has been used by companies like Meta and Microsoft. Thomas graduated from Columbia University with a B.Sc. in Artificial Intelligence.

Madison Faulkner

Madison joined NEA in 2024 and is currently a Venture Advisor on the Technology Investing Team focused on early-stage data, infrastructure, developer tools, data science, and AI/ML. Previously, she was a Vice President at Costanoa Ventures where she worked closely with Delphina.ai, Probabl.ai, Mindtrip.ai, Noteable.io (acq by Confluent), Rafay.co, and others. Prior to investing, Madison was Head of Data Science and Machine Learning at Thrasio, Head of Data Science at Greycroft, and held several data science positions at Facebook. Madison received a BS in Management Science and Engineering from Stanford University.
Madison joined NEA in 2024 and is currently a Venture Advisor on the Technology Investing Team focused on early-stage data, infrastructure, developer tools, data science, and AI/ML. Previously, she was a Vice President at Costanoa Ventures where she worked closely with Delphina.ai, Probabl.ai, Mindtrip.ai, Noteable.io (acq by Confluent), Rafay.co, and others. Prior to investing, Madison was Head of Data Science and Machine Learning at Thrasio, Head of Data Science at Greycroft, and held several data science positions at Facebook. Madison received a BS in Management Science and Engineering from Stanford University.

Lila Tretikov

Lila is Partner, Head of AI Strategy, investing across multiple stages, sectors, and geographies. She most recently served as Deputy CTO at Microsoft, driving large-scale AI transformation. Previously, Lila was SVP at Engie, CEO & Vice Chair of Terrawatt at Engie, where she led the company’s transition to new energy sources and technologies. She also served as CEO at Wikimedia Foundation, where she launched the Wikipedia Endowment, introduced an AI strategy, and reversed Wikipedia’s decline. Lila has been named one of Forbes’ Top 100 Most Powerful Women, a World Economic Forum Young Global Leader, an NACD 100 Board Director, and a distinguished alumna of the University of California, Berkeley.
Lila is Partner, Head of AI Strategy, investing across multiple stages, sectors, and geographies. She most recently served as Deputy CTO at Microsoft, driving large-scale AI transformation. Previously, Lila was SVP at Engie, CEO & Vice Chair of Terrawatt at Engie, where she led the company’s transition to new energy sources and technologies. She also served as CEO at Wikimedia Foundation, where she launched the Wikipedia Endowment, introduced an AI strategy, and reversed Wikipedia’s decline. Lila has been named one of Forbes’ Top 100 Most Powerful Women, a World Economic Forum Young Global Leader, an NACD 100 Board Director, and a distinguished alumna of the University of California, Berkeley.

Andrew Schoen

Andrew joined NEA in 2014 and is a Partner on the Technology Investing Team. His investment focus includes AI, cybersecurity, fintech, and frontier tech spanning both bits & atoms. Prior to NEA, he was a member of Blackstone’s M&A team. Prior to Blackstone, he co-founded Flicstart (acquired). Andrew serves on the Cornell University Council, the Advisory Council for Entrepreneurship at Cornell, and is President Emeritus of the Cornell Venture Capital Club. He earned his master’s degree as a Schwarzman Scholar and his bachelor’s degree in economics and engineering, Phi Beta Kappa, at Cornell University.
Andrew joined NEA in 2014 and is a Partner on the Technology Investing Team. His investment focus includes AI, cybersecurity, fintech, and frontier tech spanning both bits & atoms. Prior to NEA, he was a member of Blackstone’s M&A team. Prior to Blackstone, he co-founded Flicstart (acquired). Andrew serves on the Cornell University Council, the Advisory Council for Entrepreneurship at Cornell, and is President Emeritus of the Cornell Venture Capital Club. He earned his master’s degree as a Schwarzman Scholar and his bachelor’s degree in economics and engineering, Phi Beta Kappa, at Cornell University.