What would AI-native pharmacovigilance actually look like?
Case processing is nearly solved. The data feeding it is late and biased, and that is the part worth an AI budget.
Almost every AI programme in pharmacovigilance is aimed at case processing: reading an incoming report, coding it, spotting duplicates, drafting a narrative. That work is real and it is largely won. But case processing was never the binding constraint. A report only exists once someone chose to write one, and the median under-reporting rate is 94%. Below are seven rules the published record supports, two of which the arrival of reasoning models has just reopened.
Key takeaways
- Case processing is largely solved: France has run machine-coded patient ADR reports nationally since January 2021.
- Median under-reporting of adverse drug reactions is 94% across 37 studies in 12 countries.
- At the working threshold of EB05 = 2, signal-detection sensitivity is 0.36 against the OMOP gold standard.
- Rofecoxib's heart-attack signal was visible in European health records in 2000 and in VigiBase in Q4 2004.
- 93% of rofecoxib heart attacks reported to VigiBase had occurred before withdrawal, and were reported after it.
- The verdict that models cannot judge seriousness was measured on gradient boosting and on 1 to 8 billion parameter models.
- A reasoning model reading whole clinical notes classified seriousness correctly in 100% of 191 adverse events.
- FDA judged ARIA insufficient for 134 safety concerns, and says codes often cannot describe the clinical concept at all.
The seven rules, and where each one comes from
Each rule below links to the evidence for it further down. Two of them changed in 2026, because the studies that produced the original verdict were run on models that have since been superseded.
- 1Automate extraction now, where a labelled corpus already exists.Coding, triage, de-duplication and translation are settled work when a human stays accountable downstream. France has run this nationally since January 2021 at an AUC of 0.97.What case processing already solved
- 2Re-test the judgment tasks; the verdict against them predates the models.The seriousness F-measure of 0.63 came from gradient boosting on bag-of-words features, and the 64% causality agreement came from 1 to 8 billion parameter models. A reasoning model reading whole notes classified seriousness correctly in 100% of 191 events.The judgment boundary, and why it just moved
- 3Keep a named human on the causality call regardless of model quality.The qualified person responsible for pharmacovigilance is defined in EU law as a natural person. A better model changes their workload, not their existence.The judgment boundary, and why it just moved
- 4Pin the version and control the seed before you quote the accuracy.A validated system has to stay in a state of control. GPT-3.5 and GPT-4 failed to produce reproducible results in a GxP-oriented evaluation, which is a live tension with rule two.Why the seed matters more than the score
- 5Expect months from better statistics, not years.AbbVie's pilot caught one confirmed signal six months early at a positive predictive value of 33.3%, and confirmed nothing new on a second drug. The baseline it improves on has a sensitivity of 0.36.Better statistics buy months, not years
- 6Do not buy more spontaneous-shaped text.Social media signal detection scored an area under the curve of 0.47 to 0.53 against 0.64 to 0.69 for VigiBase. The population is still people who chose to post.More text of the same kind is not more evidence
- 7Spend the budget on the substrate, because that is where the years are.Rofecoxib was visible in European health records four years before the spontaneous database flagged it. FDA has named the blockers on active data and has been attacking the text half of them since 2021.The substrate problem was a language problem
AI-native pharmacovigilance
A drug-safety operation in which models read data generated whether or not anyone chose to report it, rather than only accelerating the handling of voluntary reports. The distinction is not model sophistication but data provenance: passive spontaneous reports carry the latency, active clinical records do not.
What case processing already solved
France has been running machine-coded patient adverse drug reaction reports nationally since January 2021. Synapse Medicine, the Bordeaux university hospital pharmacovigilance centre and the French agency took 11,633 patient-filed reports coded by 27 pharmacovigilance centres between March 2017 and December 2020, and trained models to pre-code them from the patients’ own free text. External validation on held-out centres held up, and the system has been deployed by national health authorities in France since January 2021 to support pharmacovigilance of COVID-19 vaccines. The UK regulator separately trained its own free-text coder on more than 100,000 previous vaccine reports.
The profile of a solvable job is visible in that description: a human-labelled corpus in the tens of thousands, a supervised task, and a person who stays accountable downstream. Where those three hold, this is settled work and classical methods are enough. Nothing about reasoning models changes rule one, except that they are a more expensive way to do something gradient boosting already does well.
The judgment boundary, and why it just moved
The same French paper reported a second task, and the gap between the two is what the field has been quoting ever since.
| Task | AUC | F-measure |
|---|---|---|
| Identify the adverse drug reaction | 0.97 | 0.80 |
| Assess whether it is serious | 0.85 | 0.63 |
Martin et al., Drug Safety 2022, internal validation on 11,633 patient reports from the French national portal. Best configuration per task: term frequency-inverse document frequency with gradient boosting for identification, FastText with gradient boosting for seriousness. Both are classical machine-learning pipelines, not language models.
Read as a statement about AI, that table says extraction is solved and judgment is not. Read as a statement about 2022 methods, it says something narrower: bag-of-words features and a gradient-boosted tree can find the reaction in a patient’s free text but cannot reliably weigh whether it was serious. Those are different claims, and only the second one is what was actually measured.
The same caution applies to causality. The most-cited negative result tested biomedical large language models against two human expert evaluators on 150 individual case safety reports, and the best configuration agreed with them on the final classification 64% of the time, which the authors called suboptimal for reliable causality assessment. The models tested were TinyLlama at 1.1 billion parameters, Medicine LLaMA-3 at 8 billion, and MedLLaMA. That is a fair test of small domain-tuned models. It is not a test of frontier reasoning models, and as of this writing no published study appears to have run one on formal causality assessment.
What a reasoning model did to the seriousness task
In June 2026 a group at the University of California, San Francisco published a pipeline built on OpenAI o1 that reads entire electronic health record notes rather than a single report form. Across 372 deidentified notes it surfaced 191 adverse drug events, of which 180 were true positives: 94.2% precision, 84.1% recall, F1 88.9%. On the properties that the 2022 classical pipeline struggled with, seriousness was correct in 100% of cases and label status in 93.9%, with MedDRA lowest-level-term coding correct in 92.5%. It ran at 18 cents per note and 35 cents per validated event, and it inferred events the physicians had not written down, such as tacrolimus-associated hypomagnesemia.
Hold that result to its size before rebuilding a department around it. This is 191 events at a single academic centre, with one expert reviewing validity and a 100-event gold standard behind the recall figure. The 100% on seriousness is measured on events the model itself surfaced, not on every event present in the notes, so it is not a like-for-like replacement of an F-measure of 0.63. What it does establish is that the ceiling people have been quoting was a property of the method, not of the task, and that the experiment worth running now is the comparison nobody has published.
Rule three survives all of it. Better judgment does not relocate accountability: in the European Union the qualified person responsible for pharmacovigilance is a natural person, accountable for the correctness and completeness of what is submitted. A model that is right more often changes how much work that person does, not whether they exist.
Why the seed matters more than the score
A validated system has to stay in a state of control, which means the same input has to yield a defensible output. Bayer tested GPT-3.5 and GPT-4 on pharmacovigilance entity recognition and found that both failed to demonstrate reproducible results, concluding that this, together with the limits of externally hosted systems, may affect the use of closed and proprietary models in regulated environments. A published comment on that paper followed in 2025, which is worth reading alongside it.
This is the constraint that decides architecture, and it cuts against the result in the previous section. The most capable model for the judgment tasks is currently a hosted one you cannot pin. Anyone planning to put a reasoning model inside a regulated workflow has to answer that before quoting its accuracy.
Better statistics buy months, not years
The disproportionality layer has plateaued. AbbVie published the honest ceiling: a gradient-boosted model caught one confirmed signal six months earlier than human reviewers at a positive predictive value of 33.3%, and on a second drug none of the remaining candidate signals were confirmed on human review. That is better arithmetic on the same biased sample.
It is worth knowing how weak the baseline is. At an EB05 of 2, the threshold the industry actually runs on, measured sensitivity against the Observational Medical Outcomes Partnership gold standard is 0.36. Roughly two in three real signals do not fire. And in that same validation, the events the methods handled worst were those related to acute myocardial infarction, which is the signal the next section is about.
What moving to active data is worth: rofecoxib
Three drugs matter to this story, and it helps to separate them first. Rofecoxib, sold as Vioxx, and celecoxib, sold as Celebrex, are COX-2 selective anti-inflammatories, designed to spare the stomach. Naproxen is an older non-selective anti-inflammatory of the same broad family.
The cardiovascular signal for rofecoxib appears in the primary literature in November 2000. The VIGOR trial randomised 8,076 rheumatoid arthritis patients to rofecoxib or naproxen and found myocardial infarction in 0.4% of the rofecoxib arm against 0.1% on naproxen. At the time this was widely read as naproxen being protective rather than rofecoxib being harmful. Merck withdrew rofecoxib from the market in September 2004.
Four months after that withdrawal, an FDA safety officer published a Kaiser Permanente study covering 2,302,029 person-years and 8,143 serious coronary heart disease events. It compared rofecoxib against celecoxib rather than against naproxen, because celecoxib was the drug most patients would otherwise have been taking, and that comparison removed the ambiguity: rofecoxib carried an adjusted odds ratio of 1.59 across all doses and 3.58 above 25 mg per day. The same paper reported that naproxen did not protect against serious coronary heart disease, which retired the 2000 reading.
The part that matters here is the reconstruction published in 2018, which ran both surveillance systems retrospectively over the same period.
In EU-ADR, seven European healthcare databases covering roughly 30 million patients, a strong signal was identifiable by the third quarter of 2000 at a relative risk of 4.5 (95% CI 2.84 to 6.72), peaking at 4.8 in the fourth quarter. In VigiBase the EB05 threshold of 2 was not crossed until the fourth quarter of 2004, at 2.94, after the withdrawal. And about 93% of the myocardial infarctions eventually reported to VigiBase, 2,260 of 2,422, had occurred before the withdrawal and were reported after the risk communication. The spontaneous database did not detect the event so much as record the reaction to it.
Be fair to the caveats, which the authors state themselves. The method applied to the health-record data does not correct for confounding, and any real prospective deployment carries six to twelve months of data-access lag. The authors put the advantage at four years; subtract a full year of lag and it is still three. That is the size of the prize for working on the substrate rather than the inbox.
The substrate problem was a language problem, and that is the point
VigiBase held over 40 million individual case reports as of 31 December 2024, from more than 160 countries. EudraVigilance processed 1,757,524 individual case safety reports in 2024 alone. Those curves track the reporting obligation and the reporting infrastructure. They do not track disease burden, and they are not evidence that surveillance is improving.
Meanwhile the side of the system holding the right data cannot reach it. FDA’s Sentinel covers over 1.3 billion person-years and 371 million unique patient identifiers. Its assessment of the system covering 2022 to 2024 reports that the Active Risk Identification and Analysis capability was judged insufficient for 134 safety concerns. The leading causes, in FDA’s own accounting: insufficient supplemental structured clinical data (60 concerns), inability to identify clinical concepts with available code algorithms or terminologies (43), absence of a validated code algorithm (25), and studies requiring data elements captured only in unstructured clinical notes (25).
Those counts stay abstract until you see what the concerns actually were. The most common outcome classes ARIA could not study were infections and infestations (18 concerns), nervous system disorders (10), immune system disorders (8) and cardiac disorders (8). Eight of the infection concerns were attempts to identify progressive multifocal leukoencephalopathy, and the reason ARIA could not do it is the sharpest illustration in the report: a validated claims-based algorithm exists, with a positive predictive value of 90.0%, and FDA still judged it insufficient because signal evaluation needed a more precise outcome definition, through medical chart review or an algorithm developed in the indicated population. The code was not wrong. It was not specific enough, and the specificity lived in the chart.
Those ten root causes account for 214 citations across the concerns in that table, since one concern can be blocked by more than one thing. Two of them are about clinical detail that exists in the record as prose but not as a claims code: insufficient supplemental structured clinical data at 60 citations, and data elements captured only in unstructured clinical notes at 25. Together that is 85 of 214, or roughly 40% of everything cited. FDA routes the first to EHR-claims linkage and the second explicitly to its own natural language processing work.
Be precise about which of those a model could touch. Insufficient observation time (25 concerns), a required linkage to an unavailable data source (21) and the inability to see over-the-counter medication use (9) are data-availability problems, and no language model fixes any of them. But the largest causes are about clinical detail that exists in the record as prose and not as a code. On the second-largest, FDA is blunt about the ceiling: in many cases, no program enhancement will address these limitations of using codes to describe clinical concepts. That is a regulator saying the terminology itself is the constraint.
It is worth correcting a lazy version of this argument, which we held until we read the report. FDA has not been ignoring the text problem. Since 2021 the Sentinel Innovation Center has run a programme of natural language processing projects against it, and the largest, MOSAIC-NLP, ran named entity recognition models over more than 17 million clinical notes covering roughly 109,000 patients across 112 health systems, using 4,200 hand-annotated notes as training data. It worked: the proportion of patients with documented neuropsychiatric events rose moderately when structured electronic health record data was added to claims, and considerably again when unstructured notes were added. Related work reported an area under the precision-recall curve of about 0.77 for suicide attempt and about 0.31 for sleep-related behaviours, which is a fair picture of how uneven this is by phenotype.
So the question is not whether anyone tried, but what changes when the method changes. The Sentinel projects annotate entities to train an extractor, which needs a gold standard per concept and inherits the coding problem it was meant to escape. The UCSF pipeline reading whole notes with a reasoning model reports 94.2% precision at 35 cents per validated event, and its authors propose exactly this use: that integrated with platforms such as Sentinel in the United States or DARWIN EU in Europe, the approach could surface rare, serious and unlabelled events for regulatory analysis. It reads whole notes without a per-concept gold standard, and it infers properties such as seriousness rather than only spotting mentions. One single-centre study of 191 events is not a mandate, and MOSAIC ran at a scale this has not been tested at. But the comparison between the two approaches, on the same notes, is the most testable open question in drug safety right now, and nobody has published it.
Regulators are not the reason this has not happened
The European Medicines Agency’s reflection paper, adopted 9 September 2024, is more permissive about pharmacovigilance than about any other stage of the medicine lifecycle, allowing a flexible approach to modelling and deployment including incremental learning for classification, severity scoring and signal detection. The condition attached is that the marketing authorisation holder remains responsible for validating, monitoring and documenting model performance.
FDA began publishing adverse-event data daily rather than quarterly on 22 August 2025, and its internal model went agency-wide on 2 June 2025 with an explicitly named use in summarising adverse events to support safety profile assessments. The EU AI Act, contrary to a good deal of conference-panel anxiety, does not list pharmacovigilance among its Annex III high-risk categories. The binding constraint on AI in drug safety is Good Pharmacovigilance Practices, not the AI Act, which is why reproducibility matters more than any accuracy benchmark.
What we would build
Not a better inbox. A pipeline that treats the spontaneous report as one input among several and spends its model capacity on the unstructured clinical record, mapping notes, discharge summaries and lab narratives into the concepts ARIA says it cannot currently identify. That is an ontology and extraction problem before it is a model problem, and it is now the highest-leverage application of language models in drug safety.
Then the boring half, which is where these programmes actually die: version-pinned, seed-fixed models with a reproducibility test in the validation suite, and a qualified person who can point at a specific model version and a specific evaluation set. Note the tension this creates with the previous paragraph, and do not paper over it. The reasoning models that produce the best extraction results are hosted, and hosted models are the ones Bayer could not reproduce. Resolving that is a procurement and architecture question, and it should be answered before the pilot rather than after it.
One honest limitation of this argument. Nobody credibly knows what the current system costs. We looked for a peer-reviewed cost-per-report figure and did not find one; the circulating numbers trace to consultancy market-sizing with no disclosed methodology. The UCSF paper’s 35 cents per validated event is the closest thing to a real unit cost in this post, and it covers only inference. So the business case has to be built on latency rather than on a cost-per-case comparison. That the industry cannot price the system it is trying to automate is, itself, a finding.
The extraction layer is the work, and somebody has to build it
Turning clinical notes into concepts a safety system can query is an ontology and engineering problem before it is a model problem. That is the work we do with companies: where your data creates leverage, what to build first, and custom agents that run your science the way your team actually runs it, from proof-of-concept to production.
Glossary
| Term | What it means |
|---|---|
| ICSR | Individual case safety report, the unit of exchange in spontaneous pharmacovigilance. |
| MedDRA | Medical Dictionary for Regulatory Activities, the terminology adverse events are coded into. |
| Spontaneous reporting system | A database of voluntary adverse-event reports, such as VigiBase, FAERS or EudraVigilance. It has no denominator, so incidence cannot be computed from it. |
| Disproportionality analysis | A screen asking whether a drug-event pair appears more often than the rest of the database would predict. |
| EB05 | The lower bound of the empirical Bayes geometric mean. A value above 2 is the conventional signalling threshold. |
| OMOP gold standard | A public reference set of known positive and negative drug-event pairs, used to measure how well a signal-detection algorithm performs. |
| Active surveillance | Analysis of data generated by routine care, such as claims or health records, rather than voluntary reports. |
| ARIA | Active Risk Identification and Analysis, the routine querying capability inside FDA's Sentinel system. |
| QPPV | Qualified person responsible for pharmacovigilance, defined in EU law as a natural person. |
| Reasoning model | A language model trained to work through intermediate steps before answering, such as the o-series, rather than responding in a single pass. |
| Signal latency | The interval between an adverse effect occurring in patients and a surveillance system flagging it. |
Frequently asked questions
Is AI case processing in pharmacovigilance actually working?
Yes, for extraction. France has run machine pre-coding of patient adverse drug reaction reports nationally since January 2021, at an area under the curve of 0.97 for identifying the reaction. The corpus existed, the task is supervised, and a human remains accountable downstream. That is the profile of a solvable job.
Why does automating case processing not speed up signal detection?
Because the delay is upstream of the processing. A report only exists once a person decides to write one, and the median under-reporting rate is 94%. Coding that report in two minutes rather than twenty saves cost and headcount, but it does not change when the evidence arrived.
Can a language model assess seriousness or causality?
The published negative results used classical machine learning and small biomedical models of 1 to 8 billion parameters. A 2026 pipeline built on a reasoning model classified seriousness correctly in 100% of 191 adverse events it surfaced from clinical notes. No published study has yet run a frontier reasoning model on formal causality assessment.
What is the reproducibility problem with hosted models under GxP?
A validated system has to stay in a state of control, which means the same input yields a defensible output. Bayer tested GPT-3.5 and GPT-4 on pharmacovigilance entity recognition and found both failed to produce reproducible results, concluding that this may limit closed, externally hosted models in regulated environments.
Does the EU AI Act regulate pharmacovigilance as high-risk?
Pharmacovigilance is not listed among the Annex III high-risk categories. The binding constraints on AI in drug safety come from Good Pharmacovigilance Practices instead: validation, an audit trail, and a named accountable person. The qualified person responsible for pharmacovigilance is defined as a natural person.
Where should an AI budget in drug safety actually go?
Into the substrate. Automating intake buys cost; better statistics on the same biased sample buy months. Reading data generated whether or not anyone reported it is the only layer measured in years, and the extraction step that blocked it now has published evidence behind it.
Sources
- 1.Martin GL, Jouganous J, Savidan R, et al. Validation of Artificial Intelligence to Support the Automatic Coding of Patient Adverse Drug Reaction Reports. Drug Safety 2022;45(5):535-548 (PMID 35579816). Co-authored with the French agency; the deployment claim is the authors’ own
- 2.MHRA: Impact of AI on the regulation of medical products, 30 April 2024. Regulator self-report, no performance metrics published
- 3.Hazell L, Shakir SAW. Under-Reporting of Adverse Drug Reactions: A Systematic Review. Drug Safety 2006;29(5):385-396 (PMID 16689555)
- 4.Harpaz R, DuMouchel W, LePendu P, et al. Performance of Pharmacovigilance Signal-Detection Algorithms for the FDA Adverse Event Reporting System. Clin Pharmacol Ther 2013;93(6):539-546 (PMID 23571771). The EB05 = 2 sensitivity of 0.36 and the weakness on myocardial infarction are full-text figures, not in the abstract
- 5.Brand JS, Gauffin O, Sartori D, et al. VigiBase: Resource Profile Update. Drug Safety 2026;49(6):613-629 (PMID 41618069). Authored by the Uppsala Monitoring Centre, which operates VigiBase
- 6.Heckmann NS, Papoutsi DG, Barbieri MA, et al. Biomedical Large Language Models and Prompt Engineering for Causality Assessment of Individual Case Safety Reports. Pharmaceutical Research, 23 May 2026 (PMID 42174348). University of Copenhagen with Novo Nordisk Safety Operations
- 7.EMA: Good pharmacovigilance practices, Module I, effective 2 July 2012
- 8.Dietrich J, Hollstein A. Performance and Reproducibility of Large Language Models in Named Entity Recognition. Drug Safety 2025;48(3):287-303, online 11 December 2024 (PMID 39661234). Bayer Pharmacovigilance Data Science; a model evaluation, not a production deployment. See also the published Comment by Tiffet et al., Drug Safety 2025;49(1):139-141
- 9.De Abreu Ferreira R, Zhong S, Moureaud C, et al. A Pilot, Predictive Surveillance Model in Pharmacovigilance Using Machine Learning Approaches. Advances in Therapy 2024;41(6):2435-2445 (PMID 38704799). AbbVie-authored; the drugs are anonymised, so the result cannot be independently reproduced
- 10.Caster O, Dietrich J, Kürzinger M-L, et al. Assessment of the Utility of Social Media for Broad-Ranging Statistical Signal Detection: Results from WEB-RADR. Drug Safety 2018;41(12):1355-1369 (PMID 30043385). Data period March 2012 to March 2015; extraction methods have improved since, so treat the figures as dated to that period
- 11.Ludwig D, Wang M, Buchanan J, Trinh T. A Large Language Model for Extracting Post-marketing Adverse Drug Events from Clinical Notes in the Electronic Health Record. Drug Safety, 12 June 2026 (PMID 42283794). OpenAI o1, two-pass workflow, 372 deidentified UCSF notes yielding 191 events. Single centre, one expert reviewing validity, and a 100-event gold standard behind the recall figure; the seriousness result is measured on events the model surfaced rather than on all events present in the notes
- 12.Bombardier C, Laine L, Reicin A, et al. Comparison of upper gastrointestinal toxicity of rofecoxib and naproxen (VIGOR). N Engl J Med 2000;343(21):1520-1528 (PMID 11087881). The journal issued an Expression of Concern in December 2005, which is not relied on here
- 13.Graham DJ, Campen D, Hui R, et al. Risk of acute myocardial infarction and sudden cardiac death with COX-2 selective and non-selective NSAIDs. Lancet 2005;365(9458):475-481 (PMID 15705456). The widely circulated excess-cases figure is not in this paper and is not used here
- 14.Patadia VK, Schuemie MJ, Coloma PM, et al. Can Electronic Health Records Databases Complement Spontaneous Reporting System Databases?. Front Pharmacol 2018;9:594 (PMID 29928230). Retrospective reconstruction; the authors note the method does not adjust for confounding and that real use carries 6 to 12 months of data-access lag
- 15.EMA: 2024 Annual Report on EudraVigilance
- 16.FDA Sentinel Initiative: An Assessment of the Sentinel System (2022 to 2024), September 2025. Database figures as of April 2024: over 1.3 billion person-years covering over 371 million unique patient identifiers. The 134 insufficiency determinations are 125 pre-approval and 9 post-approval, excluding pregnancy-focused assessments. Root-cause counts and the MOSAIC-NLP figures are from Table 5 and Section 3.2
- 17.EMA: Reflection paper on the use of artificial intelligence in the lifecycle of medicines, adopted 9 September 2024
- 18.FDA: FDA Begins Real-Time Reporting of Adverse Event Data, 22 August 2025, and the agency-wide AI tool announcement, 2 June 2025. Capability claims are FDA’s own
- 19.Regulation (EU) 2024/1689, the Artificial Intelligence Act, Annex III. A Digital Omnibus amendment affecting the high-risk application dates is not reflected here
More text of the same kind is not more evidence
The obvious alternative is to aim a larger model at more spontaneous-shaped text: social media, forums, patient chatter. It has been tested properly, and it failed. WEB-RADR, an Innovative Medicines Initiative consortium, pulled 4.3 million Twitter and Facebook posts across 75 drugs from March 2012 to March 2015 and ran standard signal-detection algorithms against two validated reference sets. Area under the curve for Twitter and Facebook came in at 0.47 to 0.53. For VigiBase over the same period it was 0.64 to 0.69. At best, social media flagged 16% of positive controls before their index date, against 33% for VigiBase.
The authors’ conclusion is unusually blunt for a consortium deliverable: broad statistical signal detection in these channels performs poorly and cannot be recommended at the expense of other pharmacovigilance activities. One caveat in fairness: the data period ends in 2015 and extraction methods have improved a great deal since. What has not changed is the underlying population, which is still people who chose to post.