AI in biotech

AI-native biotech in 2026: companies building the new stack

What does the AI-native biotech stack look like in 2026, and which companies are building each layer?

In April 2026, Bessemer Venture Partners mapped the startups building biology-native data infrastructure, and AstraZeneca’s R&D leadership published its framework for adopting AI across oncology R&D. Read together, they describe one system: a three-layer stack (data collection, workflow automation, lab automation) and a three-tier adoption curve. This guide walks through both, the companies to know in each layer, and where the two views meet.

By PharosBioUpdated

Key takeaways

  • AI-native biotech runs on three layers: biology-native data collection, workflow automation, and lab automation (Bessemer, 2026).
  • Nearly 90 percent of drugs entering clinical trials fail; purpose-built data infrastructure is the industry's answer.
  • AstraZeneca adopts AI in three tiers: scientist enablement, process augmentation, enterprise transformation (Cancer Discovery, 2026).
  • Startups build layers, pharma climbs tiers; both agree the moat is data infrastructure, not individual models.
  • For known unknowns use fit-for-purpose AI; unknown unknowns need curated multimodal data plus agentic frameworks.
  • Hydra provides the workflow automation layer: autonomous analysis connecting biological data to experimental execution.

Two views of the same shift

Two documents published within a week of each other in April 2026 describe the same shift from opposite ends of the industry. Bessemer Venture Partners, mapping startups, argues that the next generation of biotech winners will be defined by their data infrastructure, not their science alone. AstraZeneca’s R&D leadership, writing in Cancer Discovery, reaches the same conclusion from inside a top-10 pharma: “it is the unique data fabric of each organization and not the models themselves that will increasingly define its edge.”

The shared premise is uncomfortable arithmetic. Nearly 90 percent of drugs entering clinical trials fail, and R&D cost per approved therapy has doubled roughly every nine years. Between 2012 and 2022, some 200 AI-for-drug-discovery companies raised $18 billion, and by 2024 more than 350 biological AI models had been published; models alone did not fix the failure rate. What is different in 2026 is a coherent stack emerging underneath the models. This post maps that stack layer by layer (the VC view), walks through how an established pharma actually absorbs it (the AstraZeneca view), and closes with what the two frameworks agree on.

Biology-native data infrastructure

Biology-native data infrastructure is the layer of tools and platforms built specifically to collect, organize, and operationalize biological data for AI-driven drug development. Bessemer defines it by three principles: scalable multimodal datasets tied to a drug’s mechanism of action, agentic AI frameworks across R&D workflows, and lab automation closing experimental feedback loops.

The VC view: three layers of the AI-native biotech stack

Figure 1: The AI-native biotech stack, 2026
Three-layer market map of AI-native biotech: a biology-native data collection layer with companies arranged along the drug development continuum (Boltz, Isomorphic Labs, Insitro, Tahoe, Noetik, Xaira, Chai Discovery, Profluent, Cradle, Enveda, Bioptimus, Owkin, Artera, Unlearn, QuantHealth, Vivodyne and others), a workflow automation layer with Potato, Convoke, Phylo, and Hydra, and a lab automation layer with Benchling, Automata, Medra, Lila Sciences, Trilobio, Ganymede, Dash Bio, Tetsuwan Scientific, UniteLabs, Zeon Systems, Briefly Bio, and Intrepid Labs.

Layers after Bessemer Venture Partners (2026). Selection and placement are our reading of public materials; illustrative, not exhaustive.

Bessemer’s market map organizes the space into three layers. At the top, biology-native data collection: companies generating multimodal biological measurements at scale, purpose-built for models rather than repurposed from legacy assays. In the middle, workflow automation: agentic platforms that orchestrate tools and unify R&D processes. At the bottom, lab automation: the physical infrastructure compressing experimental timelines.

One proof point Bessemer cites for the whole thesis: every protein-targeted small-molecule cancer drug the FDA approved between 2019 and 2023 relied on structural data from the Protein Data Bank. Public data infrastructure already made one generation of drugs possible; the bet is that proprietary, purpose-built data infrastructure makes the next one.

Layer 1: biology-native data collection

These companies span the drug development continuum, from target discovery and molecular modeling through molecular design to biomarker discovery and clinical prediction. Sixteen worth knowing (stage placement is our reading of what each company does, verified against public materials in August 2026):

CompanyStageWhat they do
BoltzDiscovery & modelingMIT spinout behind Boltz-1 and Boltz-2, open-source biomolecular structure models at AlphaFold3-level performance; Boltz-2 predicts binding affinity ~1000x faster than physics-based methods
Isomorphic LabsDiscovery & modelingAlphabet/DeepMind spinout building a unified drug design engine on AlphaFold 3; $2.1B Series B, Lilly and Novartis partnerships
InsitroDiscovery & modelingMachine learning over human genetics and cell-derived disease models to pick targets and predict clinical outcomes; Lilly, Gilead, and BMS partnerships
TahoeDiscovery & modelingMassive drug-perturbation single-cell datasets; open-sourced Tahoe-100M (100M cells, 1,200 drug treatments) into Arc Institute's Virtual Cell Atlas
NoetikDiscovery + biomarkersVirtual-cell foundation models (OCTO) trained on proprietary spatial multi-omic data from 1,000+ patient tumors; five-year GSK licensing deal
Xaira TherapeuticsDiscovery + designLaunched 2024 with $1B+ committed; pairs large-scale proprietary data generation with the RFdiffusion/RFantibody protein-design lineage from David Baker's lab
Chai DiscoveryMolecular designMultimodal generative protein models; Chai-2 reports ~16% hit rates in de novo antibody design with 20 or fewer designs per target
ProfluentMolecular designProtein language models trained on a mined CRISPR-Cas atlas; OpenCRISPR-1 is the first AI-designed gene editor to edit human DNA, published in Nature
CradleMolecular designGenerative protein engineering software that scientists tune with their own assay data; used by six of the top 25 pharma companies
EnvedaMolecular designAI on mass-spectrometry metabolomics (PRISM foundation model) to find bioactive chemistry in natural product libraries; first candidate in Phase 1
BioptimusBiomarkers & clinicalPathology and multi-omic foundation models; H-Optimus trained on 1M+ whole-slide images from 800k+ patients
OwkinBiomarkers & clinicalDigital pathology, genomics, and causal AI for biomarkers predicting treatment response and relapse risk
ArteraBiomarkers & clinicalMultimodal AI prostate-cancer test with FDA de novo authorization and NCCN guideline inclusion
UnlearnClinical predictionDigital twin generators trained on ~1M patients forecast control outcomes, cutting trial enrollment 25 to 50 percent; EMA-qualified method
QuantHealthClinical predictionSimulates clinical trials in silico at patient level; reports ~90% predictive accuracy across 600+ simulated trials, works with 12 of the top 20 pharmas
VivodynePreclinical translationRobotically grown vascularized human tissues (10,000+ per run) generating human data for efficacy and safety without animal models

Layer 2: workflow automation

Workflow automation

Workflow automation in AI-native biotech is the software layer that autonomously orchestrates the tools, databases, and analyses of R&D: taking a research direction, planning the work, executing it across systems, and validating results. It sits between data generation and physical experiments, turning both into decisions.

This is the youngest layer of the stack, and its entrants approach it from very different angles: Potato structures and automates lab protocols, Convoke runs agentic competitive overviews and ideation, Phylo builds a collaborative research workspace. Rather than catalogue a category whose differentiation is still settling, we will show the layer through the example we know best.

Hydra is PharosBio’s autonomous analysis platform. Give it a research direction and it plans the analysis, executes it across ~100 scientific databases and 200+ codified skills, and validates every result before you see it. That is the workflow layer’s job description: data in, validated decisions out. And because the layer sits directly above the physical one, Hydra’s integration skills can hand designed experiments to lab automation partners and pull the measured data back, as its Adaptyv Bio skill does for protein designs. Worked examples are in our case studies.

Layer 3: lab automation

We covered this layer in depth in our guide to lab software and, more directly, in Lab automation meets computational biology (the 8 automation levels, four buying models, and ten companies including Medra, Lila Sciences, Automata, and Trilobio). Bessemer’s map adds names on the software-and-connectivity edge of the layer:

CompanyWhat they do
BenchlingThe established anchor: cloud ELN, biological registry, sample and workflow management used by 200,000+ scientists
Ganymede"Lab-as-Code" data platform connecting instruments, LIMS/ELN, and apps into one harmonized data lake
Dash BioRobotic GLP lab returning productized bioanalysis (ELISA, qPCR, LC-MS) in days instead of months; $47.5M raised
Tetsuwan ScientificRobotic "AI scientists" that plan, run, and iterate wet-lab experiments on a science-native automation OS
UniteLabsVendor-agnostic connectivity: 100+ instruments exposed through one cloud API on open SiLA standards
Zeon SystemsAI plus off-the-shelf robotic arms so scientists specify experiments in natural language; pilots at UCSF and Stanford
Briefly BioGenerative AI that converts free-text protocols into structured, automation-ready formats to fix reproducibility
Intrepid LabsAutonomous drug-formulation experiments (Valiant platform), compressing formulation development from months to days

The pharma view: AstraZeneca's three tiers of AI adoption

The startup map shows what is for sale. The harder question is how an established R&D organization actually absorbs it. In their Cancer Discovery article, AstraZeneca’s R&D leadership (Goodwin, Barry, Weatherall, Platz, and Reis-Filho) describe AI in pharmaceutical R&D as three tiers with different goals and timelines.

Figure 2: Three tiers of AI adoption in pharma R&D
Three ascending tiers of AI adoption in pharma R&D after AstraZeneca: tier 1 individual scientist enablement and tier 2 process-enabled augmentation both answer known unknowns, while tier 3 enterprise-wide transformation uses agentic frameworks over curated, traceable, semantically integrated data to find unknown unknowns.

After Goodwin et al., AstraZeneca R&D, Cancer Discovery (2026).

Tier 1 removes repetitive work for individual scientists: automated literature review, natural-language protocol generation, and exploratory data analysis beyond fixed workflows. Tier 2 builds AI into processes so it augments domain experts continuously: systems that monitor preclinical and clinical data streams, competitive intelligence engines tracking scientific and patent activity, and decision support that integrates historical project data into go/no-go calls. Tier 3 is the ambitious one: AI-driven hypothesis generation and end-to-end optimization through interconnected agentic frameworks that combine specialized models, including omics models paired with language models.

Three nuances make the framework more than a maturity ladder. First, the epistemics: tiers 1 and 2 answer known unknowns, questions you already know to ask; only tier 3 explores unknown unknowns, and it demands curated datasets, traceability to source, semantic layers, and sandbox environments before it can be trusted. Second, the economics: AI is “durational, not permanent.” Tools will be replaced, so the authors argue for annual AI budgeting, like reagents, rather than one-off transformation projects. Third, the proof: quantitative continuous scoring of TROP2 became the first computational-pathology predictive biomarker used to select patients for antibody-drug-conjugate therapy, evidence that the upper tiers are reachable, not aspirational.

The more-plex fallacy

The more-plex fallacy is the assumption that measuring more biomarkers at higher resolution automatically produces greater insight. AstraZeneca’s R&D leadership warns that mathematical oversegmentation yields statistically detectable but functionally irrelevant patterns; for well-defined questions, smaller well-curated datasets often outperform sheer data volume.

Reading the two maps together

A venture market map and a pharma adoption framework are different instruments, but they are pointed at the same object. Five points of agreement stand out:

  • Data before models. Bessemer's first principle is mechanism-informed multimodal datasets; AstraZeneca says the data fabric, not the models, defines the edge. Both treat models as increasingly commoditized.
  • Agentic AI is the operating layer. Bessemer's agentic frameworks across R&D workflows are AstraZeneca's tier-3 mechanism: interconnected agents combining specialized models.
  • The loop must close physically. Bessemer's lab automation principle mirrors AstraZeneca's design-make-test-learn loop and lab-in-the-loop experimentation.
  • Fit for purpose beats maximal complexity. The more-plex fallacy and mechanism-of-action-informed data collection are the same warning against measuring everything and hoping.
  • Infrastructure is the moat. Both locate durable advantage in the system that generates evidence, not in any single molecule or model.

Where they differ is geometry. The map is spatial: all three layers are being built simultaneously by different companies. The tiers are temporal: an organization climbs them in sequence, and each tier raises the bar for data quality and governance. Startups sell layers; pharma buys tiers. The practical question for any R&D organization is which companies from the map serve the tier you are actually on:

AZ tierWhat it needsCompanies from the map
Tier 1: scientist enablementSelf-serve analysis over integrated data; records and protocols AI can readHydra (self-serve autonomous analysis), Benchling (structured records), Briefly Bio (structured protocols)
Tier 2: process augmentationContinuous monitoring and decision support; instruments wired into data streamsHydra (continuous portfolio monitoring), Ganymede and UniteLabs (instrument connectivity), Owkin and Artera (biomarker programs), Unlearn and QuantHealth (trial design)
Tier 3: enterprise transformationCurated multimodal data, semantic layers, agentic frameworks, closed loopsNoetik, Bioptimus, and Tahoe (curated multimodal data), Hydra with integration skills (agentic analysis reaching the bench), Medra, Lila Sciences, and Dash Bio (automated execution)
The stack in one paragraphAI-native biotech stack

The AI-native biotech stack has three layers: biology-native data collection generates model-ready biological data, workflow automation turns that data into validated decisions, and lab automation turns decisions into experiments. Hydra, PharosBio’s autonomous analysis platform, is the workflow-automation layer in practice, connecting the data layer to wet-lab execution.

Whichever tier your organization is on, the workflow layer is the entry point that pays for itself first. Tier 1 is a scientist asking Hydra an exploratory question today. Tier 2 is Hydra monitoring a portfolio continuously, as in our portfolio monitoring case study. Tier 3 is Hydra handing validated designs to automated labs and learning from what comes back.

AI-native biotech glossary (quick reference)

TermMeaning
Biology-native data infrastructureTools and platforms built specifically to collect and operationalize biological data for AI
Foundation modelLarge model trained on broad data, adaptable to many downstream tasks
Virtual cellComputational model predicting how cells respond to perturbations
Perturbation dataMeasurements of cellular responses to drugs or genetic changes
PhenomicsLarge-scale imaging-based profiling of cellular phenotypes
Spatial biologyMeasuring molecules while preserving their location in tissue
Computational pathologyAI analysis of digitized tissue slides for diagnosis and prediction
Digital twinModel forecasting an individual patient's outcome under control conditions
Synthetic control armTrial comparator built from historical and predicted patient data
In-silico trialSimulated clinical trial run computationally before enrollment
Agentic AIAI systems that plan and execute multi-step work autonomously
Workflow automationAgentic orchestration of R&D tools, data, and analyses
Semantic layerShared data vocabulary that lets models and agents consume enterprise data
Data fabricAn organization's integrated, governed data landscape
FAIR dataFindable, Accessible, Interoperable, Reusable; AZ adds T for trustworthy
Known vs unknown unknownsQuestions you know to ask vs patterns you did not know to look for
More-plex fallacyAssuming more markers at higher resolution automatically means more insight
Design-make-test-learnThe iterative loop of drug discovery: design, synthesize, assay, update
Eroom's lawThe observation that drug R&D cost per approval doubles roughly every nine years
Protein language modelModel trained on protein sequences to predict and generate new proteins

Frequently asked questions

What is biology-native data infrastructure?

Biology-native data infrastructure is the layer of tools and platforms built specifically to collect, organize, and operationalize biological data for AI-driven drug development. Bessemer defines it by three principles: scalable multimodal datasets tied to a drug's mechanism of action, agentic AI frameworks across R&D workflows, and lab automation closing experimental feedback loops.

What are the three layers of the AI-native biotech stack?

Biology-native data collection generates model-ready biological data at scale, from protein structures to spatial multi-omics and clinical outcomes. Workflow automation is the agentic software layer that orchestrates tools and analyses, turning data into validated decisions. Lab automation is the physical layer that runs experiments and closes the feedback loop.

How is big pharma adopting AI in R&D?

AstraZeneca's R&D leadership describes three tiers: individual scientist enablement (exploratory analysis without fixed workflows), process-enabled augmentation (systems that monitor data streams and support go/no-go decisions), and enterprise-wide transformation (agentic frameworks generating hypotheses over curated, traceable data). They also argue AI is durational, not permanent, and should be budgeted annually like reagents.

What is the more-plex fallacy?

The more-plex fallacy is the assumption that measuring more biomarkers at higher resolution automatically produces greater insight. AstraZeneca's R&D leadership warns that mathematical oversegmentation yields statistically detectable but functionally irrelevant patterns; for well-defined questions, smaller well-curated datasets often outperform sheer data volume.

Where does Hydra fit in the AI-native biotech stack?

Hydra is the workflow automation layer in practice. Give it a research direction and it plans, executes, and validates the analysis across ~100 scientific databases and 200+ codified skills. Through integration skills, such as its Adaptyv Bio skill, it hands designed experiments to lab automation partners and pulls validated data back for the next round.

Sources

  • Hedin, A., Jalbut, M., & Dai, G. (2026). Building biology-native data infrastructure for the AI era. Bessemer Venture Partners, Atlas (April 12, 2026). Source of the layer framework, market map, and the failure-rate, funding, and Protein Data Bank statistics.
  • Goodwin, R.J.A., Barry, S.T., Weatherall, J., Platz, S.J., & Reis-Filho, J.S. (2026). Enabling AI to Drive Innovation and Precision across Oncology R&D. Cancer Discovery. doi:10.1158/2159-8290.CD-26-0271. Source of the three-tier framework, the more-plex fallacy, and the data-fabric quote.
  • Company descriptions: our reading of company websites and public announcements, verified August 2026. Notable primary items: Profluent’s OpenCRISPR-1 (Nature), Artera’s FDA de novo authorization, Unlearn’s EMA-qualified PROCOVA method, Tahoe-100M in Arc Institute’s Virtual Cell Atlas, and Noetik’s GSK licensing agreement. Stage placements within the continuum are ours, not Bessemer’s.
  • PharosBio: Lab automation meets computational biology and Best AI tools for scientists in 2026.

The workflow layer is where the stack pays off first

Hydra turns your biological data into validated decisions today, and hands designs to automated labs when you are ready to close the loop.