Jason Bryant, general manager of AI platforms at ArisGlobal, and Samuel Wallis, head of case processing at Bristol Myers Squibb, argue that pharmacovigilance teams need a formal measure of trust in validated AI systems — a governance question regulators are now framing across the whole medicines lifecycle.
Shutterstock - Inkoly
Anyone who has worked around a validated line will recognise this scenario: even when a system performs as intended, confirmed by the release checks, someone still reviews the record, sometimes twice. The same pattern is emerging in pharmacovigilance around AI-assisted adverse event processing. A system clears a case and the quality metrics confirm it, yet a reviewer goes back over the file anyway. Precision and recall scores have climbed steadily across successive validation cycles, and still this habit persists among specialists, even where operational evidence supports greater reliance.
This hampers the technology’s ROI, not least by tying up precious human capacity that could be spent elsewhere. Until reviewers feel able to stop routinely revisiting cases a validated system has already cleared, pharmacovigilance departments pay twice for the same work: once through the AI, and again through the human hours added back on top of the time the technology was meant to save.
Accuracy was never going to be the whole answer
For some routine case-processing tasks, accuracy is no longer the primary unresolved issue. The unresolved part has to do with confidence — but not around whether a system performs well on average, so much as whether a reviewer feels they can rely on its assessment of the particular case in front of them.
Some of the hesitation traces back to a mismatch in expectations. Reviewers used to rules-based software expect identical inputs to produce identical outputs time after time. AI models, however, are probabilistic rather than rules-based, and are not necessarily designed around that same expectation — even when performing exactly as validated. A system can be highly accurate and still look unreliable when judged against that older standard, even where its actual performance clears the bar set for operational use. Adding further rules rarely closes that distance in perception. Each additional constraint narrows the range of cases a model can handle without manual intervention, trading operational value for the appearance of control — usually the point when a promising pilot will stall.
Measuring trust rather than assuming it
Case narratives highlight the difference between validated accuracy and reviewer confidence in practice. Many of the edits reviewers make to AI-generated narratives reflect house style or local convention, not clinical error, yet once logged as an edit they read as a system failure regardless of what actually changed underneath. Quality control built for a pre-AI process can’t simply be layered onto an AI-driven one unchanged if organisations want the efficiencies case processing was meant to deliver. Instead, organisations need a structured way of measuring appropriate reliance, inferred primarily from operational evidence rather than self-reported sentiment.
A formal trust coefficient could do this by combining several operational signals into a single structured view: model performance, validation evidence, reviewer behaviour, task risk and an organisation’s governance maturity. As evidence accumulates that a system is performing reliably, the coefficient would support progressively lighter oversight without loosening governance.
Agentic systems change where judgement happens, but not who is accountable
A trust coefficient needs to account for agentic AI too. Earlier tools supported individual tasks; agentic systems can now complete entire sequences of preparatory work — compiling data, checking fields, assembling evidence — before a case reaches a human reviewer, shifting specialists’ attention toward interpretation and the genuinely uncertain cases that warrant their expertise.
What matters is the difference between reviewers checking every stage of a system’s output and reviewers only stepping in where risk or clinical significance genuinely demands it. None of this changes who bears responsibility: AI might shift where the groundwork happens, but accountability for patient safety remains with qualified PV professionals.
Regulators are moving in the same direction
Proportionate, risk-based thinking is emerging in regulatory guidance too. In January 2026, the European Medicines Agency and the US Food and Drug Administration jointly published 10 guiding principles covering AI use across the medicines lifecycle, spanning nonclinical, clinical and post-marketing evidence generation, and calling for validation proportional to intended use, robust data governance and continuous performance monitoring. The Council for International Organizations of Medical Sciences reached broadly similar conclusions in its December 2025 Working Group XIV report, favouring adaptable governance frameworks over rules written for specific technologies.
The scale and complexity of modern PV work adds further impetus for change. VigiBase, the World Health Organization’s database of individual case safety reports, had passed 40 million reports by the end of 2024. Processing capacity is no longer the limiting factor; knowing when a validated system has earned less routine checking is.
Exactly how a trust coefficient would be calculated remains open to debate, and it probably ought to stay that way until more organisations have tested different approaches against real data. What would help settle the question faster is collective industry candour about the safe, phased level of oversight companies have actually achieved, so no single organisation is left working that out alone.
Regulators have already framed this as a governance question for the whole medicines lifecycle, not a single department’s problem to solve alone. That framing matters: whatever methodology any one organisation arrives at, the value of a trust coefficient will come from how consistently it can be applied and compared, not from how cleverly it is calculated in isolation.
