AI governance in pharma: where to tier up, where to move faster
Where should pharma dial AI governance up, and where should it move faster?
Trust-first is not the same as governance-everywhere. Dial governance up when a wrong output reaches a regulator, a patient or a released batch and cannot be pulled back. Move fast everywhere else. The tier is set by the consequence and reach of a wrong answer, not by the technology, not by the department, and not by how confident the vendor sounds.
Key takeaways
- Set the tier by the consequence and reach of a wrong output, not by the technology or the department.
- 58 percent of pharma technology leaders say trust-first, yet only 40 percent of pilots reach scaled deployment.
- Most organisations can write the high tier. Few genuinely release the low one, and that is where speed comes from.
- Moderna reached over 1,400 employee-built tools because low-criticality work never entered a review queue.
- Johnson and Johnson halved Installation Qualification time in 2020, two years before FDA published the CSA draft.
- Epic's sepsis model ran for years at a claimed AUC of 0.76 to 0.83; independent validation found 0.63.
- Re-classify whenever the audience changes, because tier drift is how governance fails after go-live.
The short answer
Most pharma organisations can write the high tier. Very few genuinely release the low one, because nobody is ever punished for adding a review. That second half is where the speed actually comes from. The same model can sit in both tiers on the same day, depending on where its output lands.
Move fast when all three are true
A wrong output is caught by the person who asked for it, before it goes anywhere. It reaches nobody outside that person's team. Nothing downstream of it is irreversible. Meeting summaries, literature triage, draft slides, internal search and first-pass code normally clear all three.
Slow down when any one is true
The output reaches a regulator, a patient, a released batch or the public. No human reads it before it decides something. The action it triggers cannot be pulled back. It runs on a vendor model whose validation you have not reproduced on your own data. It has quietly changed audience since you classified it. Adverse event classification, regulatory submission content, dosing decisions, batch release and promotional claims sit here.
Risk tiering
Risk tiering places each AI use case on two independent axes before it is built: how bad a wrong output would be, and how far that output travels. The position sets the oversight. It replaces a single review queue with a rule that says which work needs evidence and which work needs none.
Why tier it rather than govern everything
Fifty-eight percent of technology leaders at multinational pharma and biotech companies say they are taking a "trust-first" approach to artificial intelligence, embedding governance across the entire AI lifecycle rather than bolting it on at the end. That figure comes from ZS's 2026 CDIO Research, published 4 November 2025, based on a survey of 115 US-based technology executives at multinational pharma and biotech companies, 62 percent of them at executive level, plus twelve interviews with chief information officers.
Read alone it sounds like maturity. Read the rest of the same survey and a different picture assembles itself. Only 40 percent of AI pilots reach scaled deployment. Only 17 percent of respondents can point to measurable value in research discovery today, and only 30 percent in clinical. Meanwhile 88 percent are increasing cloud and infrastructure spend and 84 percent are increasing AI platform spend.
All figures from ZS 2026 CDIO Research (source 1), n=115 technology executives at multinational pharma and biotech companies.
Near-universal investment, a majority commitment to trust-first governance, and a pilot-to-production conversion rate that would get a drug programme killed. The comfortable explanation is that AI is hard and pharma is regulated. Our reading is that many organisations have quietly translated trust-first into governance-everywhere, and governance-everywhere is a tax paid on every single use case in order to insure against the handful that carry real risk.
ZS's own researchers put it more diplomatically than we will: "Depending on the use cases, a one-size-fits-all approach to building trust can bog down innovation where speed could be a stronger lever."
The framework: two axes, three tiers
The cleanest working version we have found is Moderna's. When the company rolled out ChatGPT Enterprise it did not route every use case through a central board. It built a criticality matrix, formalised in November 2024, with two axes: impact of failure, from low to critical, and audience, from an individual to a team to the whole company. Everything is placed on that grid before it is built, and the position determines the oversight.
| Tier | What it requires |
|---|---|
| Low criticality | Compliance with the AI code of conduct. Essentially nothing else. |
| Medium criticality | Higher standards of quality, maintenance and support. |
| High criticality | Cybersecurity review, documented design and monitoring standards, and quarterly evaluation by oversight committees. |
Moderna's framework as described in the Harvard Business School case (sources 2 and 3).
The useful move is to plot your own portfolio on it. The axes below are Moderna's; the placements are our own reading of where common pharma use cases fall, and yours will differ.
Axes after Moderna's criticality matrix (sources 2 and 3). Placements are our own. The grid deliberately omits research use cases that never leave a scientist's notebook, which sit off the bottom left corner entirely.
The axes are independent
A tool used by one person can still be high tier if its output reaches a submission. A tool used by the whole company can be low tier if it only ever drafts internal text. Collapsing the two into a single importance score is the most common way to get this wrong, and it produces the outcome where a video background generator and an adverse event classifier arrive at the same committee.
Movement across the grid is the risk
A tool built low tier for one analyst, which quietly becomes the way forty people prepare a regulatory document, has changed tier without anyone re-classifying it. Governance does not usually fail at go-live. It fails at drift, and drift has no launch date to trigger a review.
Moderna's own test case was a tool called DoseID, which recommended drug dosing for clinical trials. Chief executive Stephane Bancel raised the regulatory implications immediately. Clinical trials are tightly regulated and a dosing engine is not a video background, so it went up the tiers rather than around them. His stated concern is the sharpest one-line risk definition we have seen from a pharma chief executive: "The risk is we have a hallucination and we send something to regulators that is incorrect."
Note what that sentence does. It does not say the risk is hallucination. It says the risk is hallucination that reaches a regulator. The failure mode is defined by where the output lands, which is what lets one organisation loosen and tighten at the same time without being incoherent.
What loosening actually bought
Moderna: volume, because the queue never formed
Employees built their own tools rather than requesting them. 750 custom GPTs within two months of the rollout, with 40 percent of weekly active users having built at least one (source 4), and more than 1,400 created by December 2024 according to the Harvard Business School case (source 2). The internal chat assistant, mChat, reached more than 80 percent of employees. That volume is a governance outcome, not a governance failure: the low tier was genuinely released, so most work never entered a queue.
Johnson and Johnson: the same logic inside validation
If Moderna is the strategic version, the move to Computer Software Assurance is the operational one. Computer System Validation treats every requirement as equally deserving of scripted, documented, formally executed testing. Computer Software Assurance, which the US Food and Drug Administration finalised as guidance on 24 September 2025, replaces that with risk-based critical thinking: concentrate formal testing on what is Critical-to-Quality, and use lighter, unscripted or supplier-leveraged assurance everywhere else.
The instructive part is the date. Johnson and Johnson's technology quality team presented their results in July 2020, more than two years before FDA published the CSA draft and five before it was final. One of their stated objectives was "Don't wait for the FDA Draft Guidance to be released" , and the team was represented on the FDA-Industry CSA group that helped shape it. They were an input to the guidance, not a consumer of it.
On a manufacturing execution system pilot, 5 of 11 Installation Qualification test scripts moved from scripted to unscripted, and execution and review time fell from four weeks to two. Configuration verification shifted to 90 percent informal and 10 percent formal testing, where previously all of it was formal. Out-of-the-box vendor functionality saw a 100 percent reduction in re-testing effort by relying on supplier validation.
The sentence that matters most in their material is "Formal Testing reduced and focused on Critical to Quality requirements." Reduced and focused. The rigour did not leave the building. It got concentrated where failure has consequences. These are company-reported figures from a single pilot rather than independently audited ones.
What the high tier is for
The case for tiering collapses if you cannot show what the high tier buys. Two deployments show it from opposite directions: one where governance stopped a programme at enormous cost, and one where it never engaged at all.
MD Anderson: 62 million dollars, and worth it
In 2012, MD Anderson Cancer Center partnered with IBM to build a clinical decision support tool for oncology on the Watson platform. Five years and 62 million dollars later, MD Anderson let the contract expire. A University of Texas System audit dated November 2016 and released the following February documented procurement problems, cost overruns, delays, and serious difficulty integrating the system into clinical workflow. The line from the Journal of the National Cancer Institute news feature belongs above every AI steering committee agenda:
"Five years and $62 million later, M. D. Anderson let its contract with IBM expire before anyone used Watson on actual patients."
That is governance doing the job it exists to do, at the cost it is supposed to cost. Sixty-two million dollars is painful. It is also small against the alternative, and we know something about the alternative: IBM internal documents obtained by STAT News in July 2018 showed the company's own reviewers had identified multiple examples of unsafe and incorrect treatment recommendations, traced to training the system on a small number of synthetic, hypothetical cases rather than real-world evidence.
Epic: the inverse, and the more useful lesson
Epic's sepsis prediction model was deployed across hundreds of US hospitals on the strength of the vendor's own validation, with internal documentation reporting an area under the curve of 0.76 to 0.83. In June 2021 researchers at Michigan Medicine published an independent external validation in JAMA Internal Medicine covering 27,697 patients and 38,455 hospitalisations.
| Measure | Vendor documentation | External validation |
|---|---|---|
| Area under the curve | 0.76 to 0.83 | 0.63 (95% CI 0.62 to 0.64) |
| Sensitivity at the alert threshold | Not stated | 33 percent |
| Positive predictive value | Not stated | 12 percent |
| Sepsis cases missed | Not stated | 67 percent |
| Hospitalisations generating an alert | Not stated | 18 percent |
| Flagged cases clinicians had not already caught | Not stated | 7 percent |
A model firing alerts on nearly one in five admissions, missing two thirds of the actual sepsis, and adding almost nothing to what the clinicians already knew, running live in hundreds of hospitals for years. The governance did eventually happen. It happened in 2021, as a peer-reviewed paper, written by people at one health system who decided to check. That is not a control. It is an autopsy.
Epic is the more useful of the two cases, because MD Anderson's lesson is one everybody already agrees with. The harder lesson is that the tier is set by the consequence of a wrong output, not by who built the thing or how confident they sound. A vendor's validation is an input to your assessment. It is not a substitute for it.
Why waiting for the regulators will not close the gap
It would be convenient to blame the caution on organisational timidity. It is more structural than that. The frameworks are genuinely behind, and they are not catching up on a schedule anyone can plan around.
| Instrument | Status as of September 2026 |
|---|---|
| FDA guidance on AI supporting regulatory decision-making for drugs and biologics | Published January 2025 with a seven-step risk-based credibility framework. Still draft, twenty months on. |
| FDA Computer Software Assurance guidance | Three years from draft (September 2022) to final (September 2025), for a comparatively narrow topic. |
| EMA reflection paper on AI in the medicinal product lifecycle | Adopted September 2024. Guidance, not binding law. |
| EU AI Act high-risk obligations | Pushed back under the Digital Omnibus. Standalone Annex III moved to December 2027; high-risk AI inside regulated products to August 2028. |
| GAMP 5 Second Edition (2022) | Treats AI and machine learning at a high level in Appendix D11. ISPE extended it with a standalone GAMP Guide: Artificial Intelligence in July 2025. |
| FDA count of submissions with AI components | More than 500 to CDER, covering 2016 to 2023. Still the published figure in 2026. |
Against that, FDA-authorised AI-enabled medical devices passed 1,400 by March 2026, with 331 authorised in 2025 alone, the most in the agency's history. The technology curve and the guidance curve are not converging. If your AI strategy has a dependency on regulatory clarity arriving, it has a dependency that will not resolve, and the tiering decision is yours to make in the absence of anyone making it for you.
Where each rule comes from
The guide at the top is not a set of opinions. Each rule is there because an organisation already paid for the lesson.
| Rule | The case that priced it |
|---|---|
| Ship the low tier without review | Moderna released it and reached 750 GPTs in two months, over 1,400 by year end. The queue never formed because most work never entered it. |
| Concentrate rigour rather than removing it | Johnson and Johnson halved Installation Qualification time and moved configuration testing to 90 percent informal, while keeping formal testing on Critical-to-Quality requirements. |
| Escalate on consequence, not on department | DoseID was built inside one team and still went up the tiers, because its output touched a regulated trial. |
| Stop it at the gate when the high tier says stop | MD Anderson spent 62 million dollars and ended the Watson programme before it reached a patient. |
| Reproduce vendor validation before trusting it | Epic's sepsis model shipped on a claimed AUC of 0.76 to 0.83; independent validation found 0.63, with 33 percent sensitivity and two thirds of sepsis missed. |
| Name an owner and monitor after go-live | Epic's model did not fail at launch. It failed for years, unmonitored. ZS found 63 percent citing absent business ownership and 68 percent citing weak data quality and governance as causes of AI failure. |
| Re-classify when the audience changes | The drift case has no headline attached to it yet, which is the reason to write the rule down before it is yours. |
Four rules that make the tiering hold
Classify before you build, not at go-live, because the cheapest moment to place a use case is before anyone is invested in the answer. Re-classify whenever the audience changes. Name an owner at medium tier and above, meaning a person accountable eighteen months from now, not a committee. And actually release the low tier: if a use case clears all three fast-track tests it should not see a review board at all. That last rule is the one organisations skip, and skipping it is what turns trust-first into a tax.
Where the work actually is
Writing the two axes takes an afternoon. Classifying a live portfolio against them does not. The hard part is the hundred use cases already running, most of which were never classified at all, some of which have drifted a tier since anyone looked, and each of which now needs controls proportional to where it actually sits rather than to how nervous the organisation feels.
Three things determine whether a tiering framework survives contact with that portfolio. An intake that classifies before anyone builds, because the cheapest moment to place a use case is before someone is invested in the answer. Controls at each tier that a team can genuinely satisfy, because a medium tier nobody can clear in practice is a high tier with extra steps. A worked example of placing one workflow this way, the certificate of analysis handoff with a contract manufacturer, is in can an AI agent take over CDMO paperwork. And monitoring that catches an audience change, because drift is the failure mode with no launch date to trigger a review.
That is the work we do with companies. Our data and AI strategy engagements start with where the leverage is and what to build first, which in a regulated organisation is inseparable from deciding what needs evidence and what does not. From there it is custom AI workflows on proprietary data and bespoke development from proof-of-concept to production, built so that the evidence a tier demands is a by-product of running the work rather than a documentation project bolted on afterwards.
Our in vivo augmentation and combinatorial therapy case studies are both work that had to survive scrutiny from an industrial or academic partner, not merely produce an answer. That is the same test a high-tier use case has to pass, which is why the two questions are really one question.
What we would take from this
The 58 percent is not the finding. The finding is that trust-first has become a way of saying yes to governance in general while avoiding the specific and politically expensive act of naming which use cases do not need it. Uniform governance feels safe because nobody gets fired for adding a review, but it is not risk management. It is risk avoidance spread evenly, which spends scarce oversight capacity on video backgrounds while a sepsis model runs unwatched in three hundred hospitals.
MD Anderson spent 62 million dollars to not hurt anyone. Moderna shipped over 1,400 tools in a year. Both were right, because they were doing different things at different tiers. The organisations that will be ahead in three years are not the ones with the most governance or the least. They are the ones that decided, explicitly and in writing, where the line sits, and were then disciplined enough to leave the other side of it alone.
Which side of the line is your analysis on?
We help pharma and biotech teams classify a real AI portfolio, decide what needs evidence and what does not, and build the workflows that produce that evidence as they run.
Glossary
| Term | What it means |
|---|---|
| CSV (Computer System Validation) | The traditional approach: every requirement gets scripted, documented, formally executed testing regardless of its risk. |
| CSA (Computer Software Assurance) | FDA's risk-based successor, final September 2025. Formal testing concentrates on Critical-to-Quality; everything else gets lighter assurance. |
| Critical-to-Quality | A requirement whose failure would affect product quality, patient safety or data integrity. The thing formal testing is reserved for. |
| GxP | The family of Good Practice regulations (manufacturing, clinical, laboratory) that determine which systems are regulated at all. |
| GAMP 5 | ISPE's risk-based framework for validating computerised systems in regulated life sciences. Second Edition, 2022, plus a separate AI guide, 2025. |
| Annex III (EU AI Act) | The list of standalone high-risk AI use cases carrying the heaviest obligations, now applying from December 2027. |
| AUC / AUROC | Area under the receiver operating characteristic curve. 1.0 is perfect ranking, 0.5 is a coin flip. |
| Tier drift | A use case whose audience or consequence grows after classification, so its oversight no longer matches its risk. |
Frequently asked questions
What is risk tiering for AI governance in pharma?
Risk tiering places each AI use case on a grid of two independent axes before it is built: how bad a wrong output would be, and how far that output travels. The position sets the oversight, so a low-consequence internal tool ships without review while anything reaching a regulator, a patient or a released batch gets full validation.
Which AI use cases can pharma ship without a review board?
Those where all three fast-track tests hold: a wrong output is caught by the person who asked for it, it reaches nobody outside that team, and nothing downstream is irreversible. Meeting summaries, literature triage, internal search, draft slides and first-pass code normally qualify. Releasing this tier is what creates capacity for the high one.
Is a vendor's validation enough to deploy a clinical AI model?
No. A supplier's performance claim is an input to your assessment, never a substitute for it. Epic's sepsis prediction model was deployed across hundreds of hospitals on internal documentation reporting an AUC of 0.76 to 0.83; external validation at Michigan Medicine measured 0.63 with 33 percent sensitivity.
What is Computer Software Assurance and how does it differ from CSV?
Computer System Validation treats every requirement as equally deserving of scripted, formally executed testing. Computer Software Assurance, finalised by FDA in September 2025, applies risk-based critical thinking instead: formal testing concentrates on Critical-to-Quality requirements, while lighter unscripted or supplier-leveraged assurance covers the rest. Some manufacturers adopted it years before the guidance existed.
Should pharma wait for regulators to define AI governance?
Waiting does not resolve. FDA's AI credibility guidance for drug submissions has been draft since January 2025, the EMA reflection paper remains non-binding, and the EU AI Act's high-risk obligations were pushed to December 2027 and August 2028. The tiering decision stays with the organisation.
Sources
- 1.ZS (2025). Scaling AI in pharma and biotech: 2026 outlook from ZS's CDIO research. ZS Associates, 4 November 2025. n=115 US-based technology executives at multinational pharma and biotech companies, 62 percent at executive level, plus 12 interviews with chief information officers.
- 2.Bojinov, I., Lakhani, K. R., Hildebrandt, N., & Weber, S. (2025). Moderna: Democratizing Artificial Intelligence. Harvard Business School case 9-625-070, 22 January 2025. Paywalled original.
- 3.Moderna: Democratizing Artificial Intelligence (full text PDF). Publicly readable copy of source 2, hosted on ctoforum.org.
- 4.OpenAI. Moderna. Source of the 750 GPTs in two months, 40 percent of weekly actives and 80 percent mChat adoption figures only. The 1,400 figure comes from source 2. Vendor-published, so a cross-check rather than independent verification.
- 5.US Food and Drug Administration (2025). Computer Software Assurance for Production and Quality System Software. Final guidance, Federal Register, 24 September 2025.
- 6.Schardong, R., Vazquez, D., & Moussally, J. (2020). CSA Revolution Series, Episode 2: Johnson & Johnson CSA Journey and Case Study on MES. Medical Device Innovation Consortium and the FDA-Industry CSA team, July 2020. Figures are company-reported from a single pilot, not independently audited.
- 7.Schmidt, C. (2017). M. D. Anderson Breaks With IBM Watson, Raising Questions About Artificial Intelligence in Oncology. JNCI 109(5), djx113. News feature, not a research paper.
- 8.Ross, C., & Swetlitz, I. (2018). IBM's Watson supercomputer recommended "unsafe and incorrect" cancer treatments. STAT News, 25 July 2018.
- 9.Wong, A., Otles, E., Donnelly, J. P., Krumm, A., McCullough, J., DeTroyer-Cooley, O., Pestrue, J., Phillips, M., Konye, J., Penoza, C., Ghous, M., & Singh, K. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Intern Med 181(8), 1065-70.
- 10.US Food and Drug Administration (2025). Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products. Draft guidance, January 2025. Still draft as of September 2026.
- 11.European Medicines Agency (2024). Reflection paper on the use of artificial intelligence in the lifecycle of medicines. CHMP and CVMP, adopted 9 September 2024. Guidance, not binding law.
- 12.Council of the European Union (2026). Artificial intelligence: Council gives final green light to simplify and streamline rules. 29 June 2026. Digital Omnibus on AI, high-risk deadline changes.
- 13.ISPE (2025). GAMP Guide: Artificial Intelligence. July 2025.
- 14.US Food and Drug Administration (2025). FDA proposes framework to advance credibility of AI models used for drug and biological product submissions. Press release, 6 January 2025. Source of the "more than 500 submissions with AI components, 2016 to 2023" figure, which is CDER-specific.
- 15.US Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices. Device list, updated 4 March 2026; tallies via Reuter, E., MedTech Dive.