Digital transformation in pharma: from AI pilots to real ROI
What makes AI pilots succeed in pharma, and what does digital transformation actually look like in 2026?
Pharma has adopted AI faster than any technology in its history, and most of that adoption changes nothing: LLM subscriptions burn tokens on side tasks while concrete processes stay untouched. This post lays out the data behind the gap (Benchling, MIT, Bain), the principles that separate token burn from business impact, and the radical-collaboration model that makes transformation stick.
Key takeaways
- Most pharma AI runs as free-floating tokens: chat assistants on side tasks, not structured projects improving concrete processes.
- MIT found 95 percent of GenAI pilots deliver no P&L impact; pilots with external partners succeed 67 percent of the time versus 22 for internal-only builds.
- Benchling's 2026 report: 92 percent of biotech AI leaders use AI for non-scientific work; only 22 percent span multiple R&D workflows.
- AI pilots fail on data quality and compliance, not on missing AI talent (Benchling: 55 vs 22 percent citing each as major factors).
- Bain's playbook: pick one strategic lane, orchestrate agents into one system, and redesign roles with explicit human-in-the-loop decision rights.
- Radical collaboration beats disruption: deep partnership on concrete processes is how PharosBio runs every partner project.
The token burn problem
Walk into almost any pharma or biotech organization in 2026 and you will find artificial intelligence everywhere and agentic AI almost nowhere in particular. Enterprise LLM subscriptions are live, copilots are open in every browser tab, and the token meter is spinning. What you will rarely find is the thing the technology is now capable of: a structured project in which agents own a concrete process (a screening triage, a competitive-intelligence watch, a bioanalysis pipeline) end to end, with metrics, an owner, and validated outputs.
We call this pattern free-floating tokens, and the gap it leaves is the difference between AI that feels productive and AI that changes the P&L.
Free-floating tokens
Free-floating token consumption is the pattern where an organization buys LLM capacity and leaves usage to individual discretion: chat assistants summarizing documents, drafting emails, answering side questions. Tokens get burned and work feels faster, but no process changes hands. The opposite is a structured agentic project: scoped to one concrete workflow, measured, owned, and validated.
What the data says: adoption is broad, and shallow
The Benchling 2026 Biotech AI Report (a November 2025 survey of ~104 biotech and pharma organizations that already use AI in R&D: these are the leaders, not the laggards) puts numbers on the pattern. AI use is heavy where the work is simple, and thin where transformation would live.
Adoption and transformation figures from Benchling (2026); pilot outcomes from MIT NANDA (2025).
Behind the headline bars: 92 percent of organizations use AI in non-scientific use cases (document authoring, literature search, software engineering) versus 81 percent in scientific ones, and regular use splits 59 to 44 the same way. Benchling’s own conclusion is that “AI use cases are concentrated in simpler, low-friction, predictable workflows.” Meanwhile 39 percent use AI only for single, task-specific jobs, and just 22 percent span multiple R&D workflows across teams: a scientist uses one model to mine papers, another for a hypothesis, a third to design an experiment, “each step existing independently.” Only 26 percent use co-scientist agents and only 21 percent have adopted AI for workflow orchestration; both top the planned-adoption list for the next two years, which tells you where the leaders think the value actually is.
Why pilots fail
MIT’s NANDA initiative put the hardest number on it in The GenAI Divide: State of AI in Business 2025 (Challapally, Pease, Raskar, and Chari): across 300+ public deployments, 52 executive interviews, and 153 surveyed leaders, about 95 percent of GenAI pilots deliver no measurable P&L impact. Two findings matter more than the headline.
The failure is organizational, not technical. Benchling’s respondents rank data quality and availability (55 percent citing it as a major factor) and IP, security, and compliance (50 percent) as the top pilot killers. Insufficient internal AI talent ranks near the bottom at 22 percent. Gartner meanwhile predicts over 40 percent of agentic AI projects will be scrapped by 2027, mostly for operationalization reasons, not model quality.
Who builds with you changes the odds. In the MIT data, externally purchased or partnered tools succeed roughly twice as often as internal builds, and pilots that blend internal specialists with external expertise reach a 67 percent success rate versus 22 percent for internal-only builds. The empirically winning modality is not buying a tool or building alone; it is building together. Hold that thought.
The playbook: pick a lane, orchestrate, redesign roles
Bain’s How to Break Through Pharma’s Innovation Bottleneck with AI (written for clinical development, but the logic generalizes) argues that winners will place their bets carefully, embed change at the right organizational layers, and pick a strategic lane rather than scattering pilots:
| Lane | The bet | For whom |
|---|---|---|
| Horizontal | Commit fully to site-facing tools, AI literacy, and faster enrollment; become the sponsor of choice | Site-experience-led organizations |
| Vertical | Embed AI throughout internal clinical operations; become the tech-first employer of choice | Operations-led organizations |
| Upstream | Reimagine protocol design itself; high risk, high reward | Science-motivated innovators |
| End-to-end | A few lighthouse trials to test, learn, and scale | A measured path to enterprise-wide transformation |
Beyond the lane choice, Bain’s prescriptions map exactly onto the free-floating-token diagnosis. Orchestrate development as one AI-powered system, not standalone pilots: agents that plan and execute end-to-end workflows with escalation to humans, built on shared standards so specialized agents plug in without new silos, with testing and monitoring baked in (“reliability is not optional in regulated environments”). And redesign the operating model: human-in-the-loop is “not a single checker but a defined set of decision rights embedded in workflows.” Bain recommends a RAPID-style framework (Recommend, Agree, Perform, Input, Decide) for reviewers, model stewards, and data curators, updated SOPs with escalation paths, and one accountable AI owner per workflow. The result, in their words, is “a thinner execution layer and a stronger judgment layer.”
Steps after Bain (2026); success rates from MIT NANDA (2025).
Radical collaboration, not disruption
Here is where the how matters as much as the what. Hemant Taneja, CEO of General Catalyst, has spent years arguing that the industries that matter most (healthcare above all) will not be fixed by the classic disruption playbook. His transformative principles replace it with radical collaboration: technology companies and incumbents transforming concrete operations together, from inside. General Catalyst went as far as acquiring a health system (Summa Health) to prove the model, and Taneja’s one-liner names the destination: “the Amazon of healthcare is not a $1 trillion company but a $1 trillion ecosystem.”
Radical collaboration
Radical collaboration, a term popularized by General Catalyst’s Hemant Taneja, is an innovation model in which technology companies partner deeply with incumbent institutions to transform specific operations from within, sharing risk, data, and accountability, instead of attacking the incumbent from outside with a disruptive standalone product.
The MIT data quietly validates the philosophy: 67 percent success when internal experts and external partners build together, 22 percent when organizations go it alone. Free-floating tokens are the purest go-it-alone modality there is: a subscription, no partner, no process, no shared accountability. Structured projects run in radical-collaboration mode are the opposite, and they are the configuration in which pilots actually survive contact with the P&L.
From token burn to value production: the principles
| Principle | What it replaces | Source |
|---|---|---|
| Pick one strategic lane and name an accountable owner | Scattered pilots and unowned subscriptions | Bain |
| Start with 2-3 concrete workflows with measurable pain | "Roll out a copilot to everyone" | Bain, Benchling |
| Make data AI-ready for those workflows in a focused sprint | Waiting for perfect enterprise-wide data | Bain; Benchling (55% cite data as the pilot killer) |
| Orchestrate agents into one system on shared standards | Task tools that each exist independently | Bain; Benchling (22% span multiple workflows) |
| Define human-in-the-loop as decision rights (RAPID), not a checker | Vague "human oversight" | Bain |
| Build with an external partner in radical collaboration | Internal-only builds (22% success) | MIT, Taneja |
| Validate outputs scientifically before they touch decisions | Trusting fluent but unverified answers | PharosBio |
How we put these principles to work
PharosBio has led structured AI projects in exactly this configuration with both industrial and academic partners: augmenting an in vivo research program by interpreting preclinical findings against human clinical and omics data, and prioritizing drug-combination targets in a space of hundreds of millions of candidate pairs. Each engagement follows the principles above: one concrete process, an accountable owner on both sides, explicit human decision rights, and scientific validation of every output. Using those principles, we help companies improve ROI in the areas with the highest potential and the shortest path to value. More worked examples are in our case studies.
The difference between token burn and value production in pharma AI is structure: a structured agentic project pairs an autonomous system with a concrete process, an accountable owner, explicit human decision rights, and scientific validation of outputs. PharosBio runs its partner projects in this radical-collaboration model.
Are you burning tokens or running projects?
| Signal | Token burn | Structured project |
|---|---|---|
| Scope | "Everyone has access" | Two or three named workflows |
| Owner | Nobody, or IT by default | One accountable AI owner per workflow |
| Metric | Usage and seat counts | Cycle time, quality, and cost of the process |
| Human role | Vague oversight | RAPID decision rights written into SOPs |
| Data | Whatever the chat can reach | AI-ready data for the chosen workflows |
| Outputs | Read and forgotten | Validated, escalated, fed back into the process |
| Partner | None (a subscription) | External expertise in radical collaboration |
Glossary (quick reference)
| Term | Meaning |
|---|---|
| Agentic AI | AI systems that plan and execute multi-step work autonomously |
| Free-floating tokens | LLM capacity consumed by individuals with no process attached |
| Structured agentic project | Agents owning one concrete workflow with metrics, an owner, and validation |
| Workflow orchestration | Coordinating specialized agents and tools into one end-to-end system |
| Human-in-the-loop | Defined human decision rights inside an automated workflow |
| RAPID framework | Recommend, Agree, Perform, Input, Decide: who does what in a decision |
| Model steward | Role accountable for a model's performance, retraining, and escalation |
| AI-ready data | Connected, standardized, usable data scoped to a chosen workflow |
| Lighthouse trial | A flagship trial used to test, learn, and scale AI practices |
| GenAI divide | MIT's term for the gap between high AI adoption and low transformation |
| Pilot purgatory | The state where AI pilots multiply but never scale to production |
| Radical collaboration | Deep tech-incumbent partnership transforming operations from within |
| Co-scientist agent | An AI agent that plans, runs, and interprets scientific work |
| Enterprise AI | AI deployed against organizational processes rather than individual tasks |
| AI governance | Traceability, audit, and accountability structures around AI systems |
| Digital transformation | Redesigning processes and roles around new technology, not just adopting tools |
Frequently asked questions
What does it mean that AI is deployed as free-floating tokens?
Free-floating token consumption is the pattern where an organization buys LLM capacity and leaves usage to individual discretion: chat assistants summarizing documents, drafting emails, answering side questions. Tokens get burned and work feels faster, but no process changes hands. The opposite is a structured agentic project: one concrete workflow, measured, owned, and validated.
Why do most AI pilots fail in pharma?
MIT's GenAI Divide report found 95 percent of pilots deliver no measurable P&L impact. In Benchling's pharma-specific data, the major failure factors are data quality and availability (55 percent) and IP, security, and compliance (50 percent); missing AI talent is not (22 percent). Gartner expects over 40 percent of agentic AI projects scrapped by 2027 for operationalization reasons.
What is radical collaboration in pharma AI?
Radical collaboration, a term popularized by General Catalyst's Hemant Taneja, is an innovation model in which technology companies partner deeply with incumbent institutions to transform specific operations from within, sharing risk, data, and accountability. MIT's data supports it: externally partnered pilots succeed 67 percent of the time versus 22 percent for internal-only builds.
How do pharma teams get real value from agentic AI?
Bain's five actions compress the playbook: set an enterprise mandate and pick one strategic lane; start with two or three high-value workflows; make data AI-ready for those workflows in a focused sprint; build orchestrated agent systems with explicit human-in-the-loop decision rights; and redesign roles, SOPs, and incentives around the new workflows.
How does PharosBio apply these principles?
Through structured partner projects with industrial and academic customers (see the augmenting-in-vivo and combinatorial-therapy case studies): named processes, accountable owners on both sides, validated outputs, and escalation to human decision rights, targeting the areas with the highest ROI potential and the shortest path to value.
Sources
- Benchling (2026). 2026 Biotech AI Report. November 2025 survey, N=104 biotech and pharma organizations using AI in R&D.
- Challapally, A., Pease, C., Raskar, R., & Chari, P. (2025). The GenAI Divide: State of AI in Business 2025. MIT NANDA initiative, MIT Media Lab. Press coverage. 95% of pilots with no measurable P&L impact; ~67% vs ~22% success for externally partnered vs internal-only builds.
- Bain & Company (2026). How to Break Through Pharma’s Innovation Bottleneck with AI and the RAPID decision framework.
- Hemant Taneja on transformative principles and radical collaboration: The Generalist, Inside General Catalyst’s $1B+ Bet on Fixing Healthcare; Modern Healthcare, “There should be no Amazon of healthcare”.
- Gartner prediction via Forbes (2026): over 40% of agentic AI projects may be canceled by 2027.
- PharosBio case studies: Augmenting in vivo studies and Combinatorial therapy.
Improve ROI where the path to value is shortest
We run structured AI projects with industrial and academic partners: one concrete process, accountable owners, validated outputs.