Can an AI agent take over CDMO paperwork? Start with the check
Can an AI agent take over the certificate of analysis paperwork between a pharma company and its contract manufacturer?
Not the keystrokes, and not this year. The check, yes, and you can start this quarter. An agent that reads a certificate of analysis against the approved specification and says these two do not agree writes nothing to a validated system, so it needs no validation package. An agent that keys the result into SAP and releases the batch is a different regulatory object entirely. The binding constraint is not model capability. It is that the regulation asks how well your people do this today, and nobody has measured them.
Key takeaways
- Run the first agent read-only beside the coordinator. Its disagreements are your error baseline.
- Draft Annex 22 asks that the replaced process performs at a known level. Nobody measured the coordinator.
- Draft Annex 22 bars generative models from critical GMP use, and allows them elsewhere with a responsible human.
- Frontier models score far lower on multi-step professional workflows than on generic desktop tasks.
- Telefonica O2 could automate because it already knew what each process cost per human, to the second.
- Public Health England delayed 15,841 test results because no control reconciled records in against records out.
- 88 percent of 113 audited operational spreadsheets contained errors. That is the incumbent, not a reliable human.
How a batch clears quality control today
A planner raises a purchase order in SAP and emails the production request to the contract development and manufacturing organisation anyway, because the CDMO portal connects to nothing. Weeks later a certificate of analysis arrives as a PDF attachment. A coordinator opens it beside the approved specification, reads down twenty or thirty tested attributes, assay, water content, residual solvents, microbial limits, checks each against its acceptance criteria, works out retest and expiry dates from the manufacture date, and types the numbers into an SAP inspection lot. Somewhere in that chain sits a spreadsheet one person built and nobody validated. Then a second person signs to say the first got it right.
Every step of that is something a current computer-use model can attempt. None of it is something it is allowed to finish. That gap is where the argument lives.
Shadow mode
Shadow mode runs an agent read-only and in parallel with the person who currently does the work, so it produces opinions rather than actions. Its disagreements with the human become the error baseline, the business case, and the independent test data a validation package will later demand.
When to point an agent at it, and when not to
Point an agent at it now, when all three are true
The output is a comparison rather than an entry: result against specification, date against rule, PDF against what is already in SAP, and nothing it produces writes to a record. A false positive costs a human ten seconds of looking, and a false negative leaves you exactly where you already were. And you can count the volume and the time, because an automation you cannot size is an automation you cannot justify.
Do not let it sign, when any one is true
The step ends in a release decision or a goods receipt, which US regulation requires a named person to perform and a second named person to review. Or you cannot state, as a number, how well the humans currently do the job, in which case you have nothing to hold the model to. The generative question is more permissive than it is usually reported: draft Annex 22 puts generative models and large language models outside its own scope and says they should not be used in critical GMP applications, then expressly allows them in non-critical ones provided a qualified human is responsible for the output. That is the sentence the shadow reader lives in. Annex 22 is a consultation draft rather than law, and the wider European timetable has already moved once, which we covered in what still applies after the EU AI Act delay. Waiting for the text to settle is not a plan, because the measurement it will ask for takes a quarter to produce whenever you start.
Change the organisation, not only the tool
The people who would build these workflows already build them, in spreadsheets, unsupervised. Hand them a better tool without a review discipline and you industrialise the same problem. Train the operators and then monitor the operators, because a model feeding a human decision makes that human part of the system. And say what happens to the jobs before the pilot rather than after it.
Three places an agent can sit in the CDMO loop
These are usually drawn as a maturity curve, as though a company graduates from one to the next. They are better understood as three different regulatory objects, each with its own evidence requirement, and there is no rule that says you must pass through all three.
Regulatory status is our reading of the Annex 22 consultation draft, which is not law. The draft keeps generative models out of critical GMP use rather than banning them outright, so watch and check remain open to a language model provided a qualified human owns the output. The three positions are a framing, not a standard.
Almost every pilot described in public starts at do, because that is where the headcount is. Watch is where the evidence is. Check is where the money is.
Why the ordering matters
Errors cascade. For a task of n steps at error rate e, the chance of at least one error is one minus (1 minus e) to the power n, a formula Panko adapts from Lorge and Solomon, and humans make undetected errors in about 0.5 percent of simple mechanical actions such as typing. Across the twenty-five or so fields of a typical certificate that is about a twelve percent chance any given transcription carries an uncaught error. That figure is our own arithmetic on a published error rate, not a measurement, and it is an estimate precisely because the measurement does not exist. An agent in watch mode replaces our arithmetic with your data inside a quarter.
Why detection beats execution today
The obvious alternative is to point the best available model at the keystrokes and let it drive SAP. Published benchmarks argue against it. On OSWorld, which is generic desktop work, the current frontier model scores 72.6 percent against 70.2 percent for the nearest competitor. On AutomationBench, which measures multi-step professional workflows, the same model scores 41.4 percent, with the next model at 31.4 percent and the previous flagship at 18.1 percent. The year-on-year improvement is steep. The absolute number is nowhere near a release decision. These are vendor-run evaluations on vendor-chosen tasks, and the vendor post itself blocks automated retrieval, so we take them from the published system card and read them generously.
Read the vendor evaluations generously and the conclusion still holds: autonomous execution of a long professional workflow is roughly a coin flip. Detection is a different job with a different tolerance. An agent that says this assay value does not match the specification you hold can be wrong two times in three and still be worth having, because being wrong costs a glance. We have written separately about where agentic systems fail in practice and why the failure modes cluster in retrieval and handoff rather than in reasoning.
Telefonica O2: what a measured baseline made possible
O2 in the United Kingdom is the most carefully documented back-office automation in the academic literature, and what made it work was not the software. By April 2015 it had automated 15 core processes, deployed more than 160 software robots and was handling 400,000 to 500,000 transactions a month, with a twelve-month payback and a three-year return on investment between 650 and 800 percent. Business-operations staff, not developers, built the automations.
Here is the part that transfers to pharma. O2 used a heuristic of roughly three full-time equivalents saved before it bothered automating a process, and it could apply that rule because its work-allocation system already told it what a human took, per process. The manager who ran it put it plainly: "I know to the second how long that process has taken to complete over a number of years." The screening heuristic fell out of that measurement rather than preceding it. O2 had, in commercial form, exactly the thing Annex 22 asks for and pharma does not have.
Two details carry over. O2 IT first dismissed the whole programme as screen scraping built by self-taught staff and expected to need constant babysitting, which is a fair description of what already runs in most supply-chain departments and the reason the stigma exists. And when O2 removed humans from processes it thought were simple, it had to write common-sense rules that had never been needed, because a customer pre-ordering the same handset twice looked to the software like two orders. Judgment you did not know you were buying disappears the moment you stop buying it.
Public Health England: what happens when nothing reconciles
Between 25 September and 2 October 2020, 15,841 positive COVID-19 test results did not reach England contact tracing on time. An automated pipeline loaded laboratory results into a pre-2007 spreadsheet format capped at 65,536 rows. Rows past the ceiling were dropped with no error raised. The official statement said only that files had exceeded a maximum size; the format and the row ceiling come from BBC reporting, as does the estimate that each template held only about 1,400 cases, which has been questioned since because it implies more rows per result than anyone has explained. The cases were transferred once the fault was found on 3 October, so this was a delay rather than a permanent loss. No government figure was ever published for contacts missed. The widely repeated 48,000 came from the shadow health secretary extrapolating in the Commons on 5 October, and should be read as his estimate rather than as a count.
There was no model and no agent here, only an automation nobody owned, validated or monitored, which is the artefact already sitting inside most supply-chain departments. Note what actually failed. Not the arithmetic: the pipeline did what it was built to do. What was missing was any check that the records going in matched the records coming out. It failed silently because nothing was watching. An agent whose only job is to reconcile two representations of the same data is precisely the control that was absent.
Why the gap exists: outsourcing, warning letters and unvalidated spreadsheets
Pharma is not careless. It is exhaustive about the things it has decided to be exhaustive about, and blind everywhere else. Most new drugs now depend on this interface. Of the novel drugs FDA approved in 2025, 65 percent used outsourced finished-dose manufacturing, an eleven-year high, and 73 percent used outsourced active-ingredient manufacturing, just short of the 74 percent recorded in 2024. Every one of those relationships runs on documents crossing a company boundary.
The regulator already knows the documents are the weak point. Of the 303 drug and biologics warning letters FDA issued in fiscal 2025, 48 cited 21 CFR 211.84(d)(1) or (d)(2), the third most-cited current good manufacturing practice regulation that year. Those clauses are stricter than they are usually remembered: a manufacturer may lean on a supplier report of analysis only if it still runs at least one specific identity test itself and periodically validates that the supplier results are reliable. And the tooling in the gap is what you would fear. In a 2022 letter FDA recorded that a firm calculated active-ingredient assay results for validation lots on a non-validated spreadsheet whose formulas were printed at the time of calculation and never saved, so no electronic copy could be produced during the inspection.
Set that beside what the spreadsheet literature has known for two decades. Across the seven studies since 1995 that reported how many spreadsheets they audited, 88 percent of the 113 examined contained errors. Individual reviewers find roughly half the errors seeded into a document. And in one experiment the people who had just built a spreadsheet put the chance it contained an error at 18 percent on average, against an actual 86 percent. They were business students rather than professional developers, which makes the gap easier to dismiss and no less uncomfortable.
So the honest comparison is not a validated agent against a reliable human. It is a measured agent against an unmeasured spreadsheet, run by someone confident and wrong about how often it fails. Annex 22 does not obstruct that comparison. It demands somebody finally run it. The same logic drives how to tier AI governance by consequence, where the low tier exists precisely so that scarce oversight lands on the work that can hurt someone.
Where each rule comes from
| Rule | What it rests on |
|---|---|
| Start with the check, not the keystroke | Benchmark completion on multi-step professional workflows sits far below generic desktop task completion. |
| Measure the human before replacing them | Draft Annex 22 requires the replaced process to perform at a known level. |
| A measured baseline is what makes automation possible | O2 knew each process to the second, which is what let it set a three-FTE threshold. |
| Do not let the agent hold the pen | 21 CFR 211.194 requires a named performer and a named second reviewer. |
| Reconciliation is the missing control | Public Health England dropped 15,841 records with no error raised. |
| The shadow automation is already there | FDA recorded a non-validated spreadsheet used for assay calculations, formulas never saved. |
| Assume the builders are overconfident | Spreadsheet developers estimated their error likelihood far below the audited rate. |
| Removing the human removes judgment you did not price | O2 had to add common-sense rules only after automating. |
| The interface is worth fixing | Most recent novel approvals depend on outsourced manufacturing and the documents that cross that boundary. |
How to run the shadow reader
Build the shadow reader first. It needs a mailbox, a PDF parser, the specification already sitting in your ERP, and an agent that says these two things do not agree. It writes nothing to a validated system, so it needs no validation package. Run it beside the coordinator for a quarter and you end up with three numbers you have never had: how many certificates carry a discrepancy, how many the coordinator caught, and how many the agent caught that the coordinator did not.
Only then do the return-on-investment arithmetic, and do it the way O2 did: transactions per week, minutes per transaction, error rate, and a threshold below which you do not bother. We have deliberately not given a headline hours-saved figure, because every number we could find for certificate processing time traces back to a vendor selling certificate processing software. That the industry cannot say what this costs is not a gap in the argument. It is the argument.
The uncomfortable part, and our reading rather than anyone official position, is that pharma caution about AI in manufacturing has never really been caution about error rates. It has been caution about attributable error rates. An unvalidated spreadsheet fails invisibly and nobody is named. A model fails visibly, with a log, a confidence score and a version number. Annex 22 asks for exactly that. It reads as a burden because it is a standard the incumbent process has never once had to meet. The companion question, how you qualify a system that does not give the same answer twice, is the subject of validating a non-deterministic agent under GxP.
What this takes to actually run
Shadow mode is cheap to describe and awkward to staff. Somebody has to sit the agent beside a real coordinator, agree what counts as a disagreement before the first one arrives, keep the log in a form a quality function will later accept as test data, and resist the pull to let it start writing the moment it looks good. That is a measurement exercise wearing the clothes of a software project, and it fails for organisational reasons far more often than technical ones.
That is the work we do with companies. Our data and AI strategy engagements start with where the leverage is and what to build first, which in a regulated supply chain means deciding which handoffs are worth measuring before anything is automated. From there it is custom AI workflows on proprietary data and bespoke development from proof-of-concept to production. Our in vivo augmentation and portfolio monitoring case studies are both work that had to survive scrutiny from an industrial or academic partner rather than simply produce an answer, which is the same test a quality function will apply here.
The pattern is not specific to incoming goods. Any regulated process built on documents crossing a boundary has the same shape: a high-volume clerical layer nobody has measured, a signature that has to stay with a person, and a case for automation that cannot be made until the baseline exists. We applied the same reasoning to adverse event intake in what AI-native pharmacovigilance would actually look like, and the constraint there was identical.
Do you know what your CoA review actually costs?
Almost nobody does, which is why the business case never survives contact with quality. If you are looking to automate a supply chain handoff, or any regulated process, in a way that will still stand up when somebody audits it, we measure the handoff first and automate only the part that measurement justifies.
Glossary
| Term | What it means |
|---|---|
| CDMO | Contract development and manufacturing organisation: the third party that makes the drug substance or product on your behalf. |
| CoA (certificate of analysis) | The document a manufacturer issues for a batch, listing each tested attribute, the result, and the acceptance criteria it was judged against. |
| Specification | The approved list of attributes and acceptance criteria a material must meet. The CoA is checked against it. |
| Inspection lot | The SAP object that records incoming-goods quality inspection for a batch and gates the goods receipt. |
| Annex 22 | The EudraLex Volume 4 annex on artificial intelligence, in consultation draft. Not law at the time of writing. |
| Second-person review | The independent check by someone other than the performer, required for laboratory records under 21 CFR 211.194. |
| Shadow mode | Running an agent read-only and in parallel with the human, so it generates opinions and a measurable disagreement rate rather than actions. |
| Straight-through processing | A transaction completing end to end with no human touch. The destination for some CoA traffic, and the wrong place to start. |
Frequently asked questions
Can an AI agent process a certificate of analysis today?
It can read one and compare it against a specification today. It cannot key the result into a validated system or trigger a goods receipt, because that step ends in a record someone must sign. Start with comparison, which writes nothing, and keep the human holding the pen.
What does draft Annex 22 require before automating a GMP task?
Among other things, that the model performs at least as well as the process it replaces. The trap is the second half: you have to know how well that process performs. Most companies have never measured the person doing the job, so the baseline has to be built before the business case.
What is shadow mode for an AI agent?
Running the agent read-only and in parallel with the person who currently does the work, so it produces opinions rather than actions. Its disagreements with the human become the error baseline, the business case, and the independent test data that a validation package will later ask for.
Why not point a computer-use agent straight at SAP?
Because published benchmarks put multi-step professional workflow completion far below generic desktop task completion, which is nowhere near a release decision. Detection tolerates that error rate and execution does not. A wrong flag costs a glance; a wrong entry enters a regulated record.
Who signs when an agent reviews a batch record?
A person. US regulation requires the initials of whoever performed a test and of the second person who reviewed the record. An agent has no initials and no training file, so it can prepare and flag, but the attributable signature stays with a named human.
Sources
- 1.European Commission (2025). EudraLex Volume 4, Annex 22: Artificial Intelligence. Draft for consultation, 7 July 2025; consultation closed 7 October 2025. Six pages, drafted by the EMA GMDP Inspectors Working Group with PIC/S, FDA and MHRA as observers. Clauses referred to here are section 1 (scope), section 4.3 (no decrease in performance), sections 3.3 and 10.5 (human-in-the-loop), section 9 (confidence and thresholds) and sections 10.3 and 10.4 (monitoring in operation). This is a draft, not law. No final text had been adopted at the time of writing, and an EMA workshop on 30 June and 1 July 2026 considered whether safeguards could bring generative models further into scope.
- 2.Lacity, M. C., & Willcocks, L. P. (2016). Robotic Process Automation at Telefonica O2. MIS Quarterly Executive 15(1), 21-35. Peer-reviewed case study built on interviews with O2 and its software vendor; the return-on-investment figures are the company's own and were not independently audited. The April 2015 cut-off comes from the earlier LSE working paper by Lacity, Willcocks and Craig, not from the journal article.
- 3.US Food and Drug Administration. 21 CFR 211.194(a)(7) and (a)(8), Laboratory records. The regulation asks for the initials or signature of the person performing each test and the dates, and the initials or signature of a second person showing the original records have been reviewed for accuracy, completeness and compliance.
- 4.Panko, R. R. (1998, revised May 2008). What We Know About Spreadsheet Errors. Journal of End User Computing 10(2), 15-21, doi:10.4018/joeuc.1998040102. The May 2008 revision reports 88 percent of 113 spreadsheets across the seven studies since 1995 that stated their sample size, the 0.5 percent mechanical and 5 percent logical error rates, and the Panko and Halverson (2001) experiment in which subjects estimated 18 percent against an actual 86 percent. Panko's canonical university URL no longer resolves and the paper has not been migrated, so the link above goes to his current field-audit tables, where the same data is now reported as 91 percent of 55 or 83 percent of 163 depending on the subset. The 0.5 percent figure is a real quote from the 2008 text whose framing he has since widened to a range of 1 to 5 percent. The cascade equation is adapted from Lorge and Solomon (1955). The 11.8 percent figure in this article is our own arithmetic on that equation, not a measured result.
- 5.OpenAI (2026). GPT-6 Astra, 3 September 2026, with benchmark detail taken from the accompanying system card because the announcement blocks automated retrieval. All figures are vendor-reported and run in the vendor's own harness, with competitor scores reproduced by the vendor. OSWorld figures are for the offline subset.
- 6.Public Health England (2020). PHE statement on delayed reporting of COVID-19 cases, 4 October 2020. The 15,841 figure and the dates are official; the spreadsheet row ceiling was established by contemporaneous BBC reporting rather than by the statement, and the figure of roughly 48,000 contacts is an extrapolation offered by the shadow health secretary in the House of Commons on 5 October 2020, not an official count. No government figure for contacts missed was ever published.
- 7.Slabodkin, G. (2026). Analysis of FDA drug approvals indicates favorable longer-term trend for CDMOs. Pharma Manufacturing, 5 February 2026. Trade-press report of a sell-side analysis; we read the trade report, not the underlying note, which has no public landing page.
- 8.Hartmann, E., & Oestreich, L. (2026). Trends in FDA FY 2025 warning letters. Pharmaceutical Online, 13 March 2026. Of 303 drug and biologics warning letters in FY2025 the authors analysed 135 inspection-based letters, 48 of which cited 21 CFR 211.84(d)(1) or (d)(2). We quote the count rather than a rate, because the 135 includes clinical-investigator letters that cannot cite Part 211.
- 9.US Food and Drug Administration (2022). Warning letter to Specialty Process Labs LLC, MARCS-CMS 624281, 3 May 2022, following a November 2021 inspection. Note this is an active-ingredient manufacturer, so the letter is written against ICH Q7 and the adulteration provision rather than against 21 CFR 211.194. It is cited here as evidence of the tooling in use, not as enforcement of the laboratory-records rule.