AI in biotech

Digital transformation in pharma: from AI pilots to real ROI

What makes AI pilots succeed in pharma, and what does digital transformation actually look like in 2026?

Pharma has adopted AI faster than any technology in its history, and most of that adoption changes nothing: LLM subscriptions burn tokens on side tasks while concrete processes stay untouched. This post lays out the data behind the gap (Benchling, MIT, Bain), the principles that separate token burn from business impact, and the radical-collaboration model that makes transformation stick.

By PharosBioPublished on 9 min read

Key takeaways

  • Most pharma AI runs as free-floating tokens: chat assistants on side tasks, not structured projects improving concrete processes.
  • MIT found 95 percent of GenAI pilots deliver no P&L impact; pilots with external partners succeed 67 percent of the time versus 22 for internal-only builds.
  • Benchling's 2026 report: 92 percent of biotech AI leaders use AI for non-scientific work; only 22 percent span multiple R&D workflows.
  • AI pilots fail on data quality and compliance, not on missing AI talent (Benchling: 55 vs 22 percent citing each as major factors).
  • Bain's playbook: pick one strategic lane, orchestrate agents into one system, and redesign roles with explicit human-in-the-loop decision rights.
  • Radical collaboration beats disruption: deep partnership on concrete processes is how PharosBio runs every partner project.

The token burn problem

Walk into almost any pharma or biotech organization in 2026 and you will find artificial intelligence everywhere and agentic AI almost nowhere in particular. Enterprise LLM subscriptions are live, copilots are open in every browser tab, and the token meter is spinning. What you will rarely find is the thing the technology is now capable of: a structured project in which agents own a concrete process (a screening triage, a competitive-intelligence watch, a bioanalysis pipeline) end to end, with metrics, an owner, and validated outputs.

We call this pattern free-floating tokens, and the gap it leaves is the difference between AI that feels productive and AI that changes the P&L.

Free-floating tokens

Free-floating token consumption is the pattern where an organization buys LLM capacity and leaves usage to individual discretion: chat assistants summarizing documents, drafting emails, answering side questions. Tokens get burned and work feels faster, but no process changes hands. The opposite is a structured agentic project: scoped to one concrete workflow, measured, owned, and validated.

What the data says: adoption is broad, and shallow

The Benchling 2026 Biotech AI Report (a November 2025 survey of ~104 biotech and pharma organizations that already use AI in R&D: these are the leaders, not the laggards) puts numbers on the pattern. AI use is heavy where the work is simple, and thin where transformation would live.

Figure 1: High adoption, low transformation
Chart contrasting AI adoption and transformation in pharma R&D: 92% use AI in non-scientific work, 89% of scientists reach for a copilot first, 81% use AI in scientific use cases, and 71% have adopted assistants, yet only 22% span multiple R&D workflows, 26% use co-scientist agents, 21% adopted workflow orchestration, and 24% use AI in IND submissions. A lower band shows MIT's findings: 95% of GenAI pilots deliver no P&L impact, with 67% success for externally partnered builds versus 22% for internal-only builds.

Adoption and transformation figures from Benchling (2026); pilot outcomes from MIT NANDA (2025).

Behind the headline bars: 92 percent of organizations use AI in non-scientific use cases (document authoring, literature search, software engineering) versus 81 percent in scientific ones, and regular use splits 59 to 44 the same way. Benchling’s own conclusion is that “AI use cases are concentrated in simpler, low-friction, predictable workflows.” Meanwhile 39 percent use AI only for single, task-specific jobs, and just 22 percent span multiple R&D workflows across teams: a scientist uses one model to mine papers, another for a hypothesis, a third to design an experiment, “each step existing independently.” Only 26 percent use co-scientist agents and only 21 percent have adopted AI for workflow orchestration; both top the planned-adoption list for the next two years, which tells you where the leaders think the value actually is.

Why pilots fail

MIT’s NANDA initiative put the hardest number on it in The GenAI Divide: State of AI in Business 2025 (Challapally, Pease, Raskar, and Chari): across 300+ public deployments, 52 executive interviews, and 153 surveyed leaders, about 95 percent of GenAI pilots deliver no measurable P&L impact. Two findings matter more than the headline.

The failure is organizational, not technical. Benchling’s respondents rank data quality and availability (55 percent citing it as a major factor) and IP, security, and compliance (50 percent) as the top pilot killers. Insufficient internal AI talent ranks near the bottom at 22 percent. Gartner meanwhile predicts over 40 percent of agentic AI projects will be scrapped by 2027, mostly for operationalization reasons, not model quality.

Who builds with you changes the odds. In the MIT data, externally purchased or partnered tools succeed roughly twice as often as internal builds, and pilots that blend internal specialists with external expertise reach a 67 percent success rate versus 22 percent for internal-only builds. The empirically winning modality is not buying a tool or building alone; it is building together. Hold that thought.

The playbook: pick a lane, orchestrate, redesign roles

Bain’s How to Break Through Pharma’s Innovation Bottleneck with AI (written for clinical development, but the logic generalizes) argues that winners will place their bets carefully, embed change at the right organizational layers, and pick a strategic lane rather than scattering pilots:

LaneThe betFor whom
HorizontalCommit fully to site-facing tools, AI literacy, and faster enrollment; become the sponsor of choiceSite-experience-led organizations
VerticalEmbed AI throughout internal clinical operations; become the tech-first employer of choiceOperations-led organizations
UpstreamReimagine protocol design itself; high risk, high rewardScience-motivated innovators
End-to-endA few lighthouse trials to test, learn, and scaleA measured path to enterprise-wide transformation

Beyond the lane choice, Bain’s prescriptions map exactly onto the free-floating-token diagnosis. Orchestrate development as one AI-powered system, not standalone pilots: agents that plan and execute end-to-end workflows with escalation to humans, built on shared standards so specialized agents plug in without new silos, with testing and monitoring baked in (“reliability is not optional in regulated environments”). And redesign the operating model: human-in-the-loop is “not a single checker but a defined set of decision rights embedded in workflows.” Bain recommends a RAPID-style framework (Recommend, Agree, Perform, Input, Decide) for reviewers, model stewards, and data curators, updated SOPs with escalation paths, and one accountable AI owner per workflow. The result, in their words, is “a thinner execution layer and a stronger judgment layer.”

Figure 2: Designing an AI process in the enterprise
Five-step pipeline for designing an enterprise AI process, after Bain: pick a strategic lane and name an owner, choose two to three workflows, run a data-ready sprint, orchestrate agents with humans in the loop using RAPID decision rights, and redesign roles and incentives, with a loop that measures cycle time, quality, and cost before scaling to the next workflow, and a note that every step should run in radical collaboration with external partners.

Steps after Bain (2026); success rates from MIT NANDA (2025).

Radical collaboration, not disruption

Here is where the how matters as much as the what. Hemant Taneja, CEO of General Catalyst, has spent years arguing that the industries that matter most (healthcare above all) will not be fixed by the classic disruption playbook. His transformative principles replace it with radical collaboration: technology companies and incumbents transforming concrete operations together, from inside. General Catalyst went as far as acquiring a health system (Summa Health) to prove the model, and Taneja’s one-liner names the destination: “the Amazon of healthcare is not a $1 trillion company but a $1 trillion ecosystem.”

Radical collaboration

Radical collaboration, a term popularized by General Catalyst’s Hemant Taneja, is an innovation model in which technology companies partner deeply with incumbent institutions to transform specific operations from within, sharing risk, data, and accountability, instead of attacking the incumbent from outside with a disruptive standalone product.

The MIT data quietly validates the philosophy: 67 percent success when internal experts and external partners build together, 22 percent when organizations go it alone. Free-floating tokens are the purest go-it-alone modality there is: a subscription, no partner, no process, no shared accountability. Structured projects run in radical-collaboration mode are the opposite, and they are the configuration in which pilots actually survive contact with the P&L.

From token burn to value production: the principles

PrincipleWhat it replacesSource
Pick one strategic lane and name an accountable ownerScattered pilots and unowned subscriptionsBain
Start with 2-3 concrete workflows with measurable pain"Roll out a copilot to everyone"Bain, Benchling
Make data AI-ready for those workflows in a focused sprintWaiting for perfect enterprise-wide dataBain; Benchling (55% cite data as the pilot killer)
Orchestrate agents into one system on shared standardsTask tools that each exist independentlyBain; Benchling (22% span multiple workflows)
Define human-in-the-loop as decision rights (RAPID), not a checkerVague "human oversight"Bain
Build with an external partner in radical collaborationInternal-only builds (22% success)MIT, Taneja
Validate outputs scientifically before they touch decisionsTrusting fluent but unverified answersPharosBio

How we put these principles to work

PharosBio has led structured AI projects in exactly this configuration with both industrial and academic partners: augmenting an in vivo research program by interpreting preclinical findings against human clinical and omics data, and prioritizing drug-combination targets in a space of hundreds of millions of candidate pairs. Each engagement follows the principles above: one concrete process, an accountable owner on both sides, explicit human decision rights, and scientific validation of every output. Using those principles, we help companies improve ROI in the areas with the highest potential and the shortest path to value. More worked examples are in our case studies.

The claimStructured agentic project

The difference between token burn and value production in pharma AI is structure: a structured agentic project pairs an autonomous system with a concrete process, an accountable owner, explicit human decision rights, and scientific validation of outputs. PharosBio runs its partner projects in this radical-collaboration model.

Are you burning tokens or running projects?

SignalToken burnStructured project
Scope"Everyone has access"Two or three named workflows
OwnerNobody, or IT by defaultOne accountable AI owner per workflow
MetricUsage and seat countsCycle time, quality, and cost of the process
Human roleVague oversightRAPID decision rights written into SOPs
DataWhatever the chat can reachAI-ready data for the chosen workflows
OutputsRead and forgottenValidated, escalated, fed back into the process
PartnerNone (a subscription)External expertise in radical collaboration

Glossary (quick reference)

TermMeaning
Agentic AIAI systems that plan and execute multi-step work autonomously
Free-floating tokensLLM capacity consumed by individuals with no process attached
Structured agentic projectAgents owning one concrete workflow with metrics, an owner, and validation
Workflow orchestrationCoordinating specialized agents and tools into one end-to-end system
Human-in-the-loopDefined human decision rights inside an automated workflow
RAPID frameworkRecommend, Agree, Perform, Input, Decide: who does what in a decision
Model stewardRole accountable for a model's performance, retraining, and escalation
AI-ready dataConnected, standardized, usable data scoped to a chosen workflow
Lighthouse trialA flagship trial used to test, learn, and scale AI practices
GenAI divideMIT's term for the gap between high AI adoption and low transformation
Pilot purgatoryThe state where AI pilots multiply but never scale to production
Radical collaborationDeep tech-incumbent partnership transforming operations from within
Co-scientist agentAn AI agent that plans, runs, and interprets scientific work
Enterprise AIAI deployed against organizational processes rather than individual tasks
AI governanceTraceability, audit, and accountability structures around AI systems
Digital transformationRedesigning processes and roles around new technology, not just adopting tools

Frequently asked questions

What does it mean that AI is deployed as free-floating tokens?

Free-floating token consumption is the pattern where an organization buys LLM capacity and leaves usage to individual discretion: chat assistants summarizing documents, drafting emails, answering side questions. Tokens get burned and work feels faster, but no process changes hands. The opposite is a structured agentic project: one concrete workflow, measured, owned, and validated.

Why do most AI pilots fail in pharma?

MIT's GenAI Divide report found 95 percent of pilots deliver no measurable P&L impact. In Benchling's pharma-specific data, the major failure factors are data quality and availability (55 percent) and IP, security, and compliance (50 percent); missing AI talent is not (22 percent). Gartner expects over 40 percent of agentic AI projects scrapped by 2027 for operationalization reasons.

What is radical collaboration in pharma AI?

Radical collaboration, a term popularized by General Catalyst's Hemant Taneja, is an innovation model in which technology companies partner deeply with incumbent institutions to transform specific operations from within, sharing risk, data, and accountability. MIT's data supports it: externally partnered pilots succeed 67 percent of the time versus 22 percent for internal-only builds.

How do pharma teams get real value from agentic AI?

Bain's five actions compress the playbook: set an enterprise mandate and pick one strategic lane; start with two or three high-value workflows; make data AI-ready for those workflows in a focused sprint; build orchestrated agent systems with explicit human-in-the-loop decision rights; and redesign roles, SOPs, and incentives around the new workflows.

How does PharosBio apply these principles?

Through structured partner projects with industrial and academic customers (see the augmenting-in-vivo and combinatorial-therapy case studies): named processes, accountable owners on both sides, validated outputs, and escalation to human decision rights, targeting the areas with the highest ROI potential and the shortest path to value.

Sources

Improve ROI where the path to value is shortest

We run structured AI projects with industrial and academic partners: one concrete process, accountable owners, validated outputs.