CANON Named Owner Principle · every AI deployment requires two named persons in the audit trail, a Governance Owner and a Decision Owner, not one substituting for the other WORKING PAPER №01 The handoff that isn’t · how clinical AI escapes accountability · Mo Johnson, MD MBA EVIDENCE Duke-Margolis 2026 · most US health systems have not named who owns the clinical AI decision when something goes wrong CANON Layer 4 · Clinical AI Governance at the bedside · the layer where the named owner has to live FRAMEWORK Clinical AI Accountability Canvas™ · the diagnostic framework distinguishing Clinical AI Governance from General AI Governance EVIDENCE Stanford MedAgentBench · agentic systems already executing clinical recommendations without a named adjudicator on the chart CANON The Two Inputs · internal data the institution audits · external data the model was trained on, rarely audited at the deployment site CITATION MedicoVigilance™ Issue 6 · The Layer With No Name · 1,627 institutional subscribers CANON The Accountability Gap™ · the structural failure point where AI stops and the named owner starts FRAMEWORK Mind the 9 Blocks™ · the nine institutional blocks that must be in place before clinical AI deployment EVIDENCE npj Digital Medicine · the four-layer governance cascade · most institutions have built the first two layers and left Layer 4 unbuilt FRAMEWORK The Accountability Gap™ (TAG™) · two seats, one handoff, any regulated sector · clinical AI and financial AI WORKING PAPER №02 The seam · how a bank's automated triage system closed the alerts nobody read · Mo Johnson, MD MBA FRAMEWORK The Accountability Canvas · nine blocks, one Gap Score™, any regulated deployment · clinical AI and financial AI PORTFOLIO GPe Research · adjudication infrastructure for regulated sectors, in clinical AI and financial AI. CANON Named Owner Principle · every AI deployment requires two named persons in the audit trail, a Governance Owner and a Decision Owner, not one substituting for the other WORKING PAPER №01 The handoff that isn’t · how clinical AI escapes accountability · Mo Johnson, MD MBA EVIDENCE Duke-Margolis 2026 · most US health systems have not named who owns the clinical AI decision when something goes wrong CANON Layer 4 · Clinical AI Governance at the bedside · the layer where the named owner has to live FRAMEWORK Clinical AI Accountability Canvas™ · the diagnostic framework distinguishing Clinical AI Governance from General AI Governance EVIDENCE Stanford MedAgentBench · agentic systems already executing clinical recommendations without a named adjudicator on the chart CANON The Two Inputs · internal data the institution audits · external data the model was trained on, rarely audited at the deployment site CITATION MedicoVigilance™ Issue 6 · The Layer With No Name · 1,627 institutional subscribers CANON The Accountability Gap™ · the structural failure point where AI stops and the named owner starts FRAMEWORK Mind the 9 Blocks™ · the nine institutional blocks that must be in place before clinical AI deployment EVIDENCE npj Digital Medicine · the four-layer governance cascade · most institutions have built the first two layers and left Layer 4 unbuilt FRAMEWORK The Accountability Gap™ (TAG™) · two seats, one handoff, any regulated sector · clinical AI and financial AI WORKING PAPER №02 The seam · how a bank's automated triage system closed the alerts nobody read · Mo Johnson, MD MBA FRAMEWORK The Accountability Canvas · nine blocks, one Gap Score™, any regulated deployment · clinical AI and financial AI PORTFOLIO GPe Research · adjudication infrastructure for regulated sectors, in clinical AI and financial AI.
POSITION PAPER №03
Abstract

Physicians override approximately 90% of drug-drug interaction alerts generated by clinical decision support systems, a figure pooled across 570,776 prescriptions in a 2024 systematic review and meta-analysis, and consistent with override rates reported since 2006. Adjusting the rules does not move the number. This paper argues that the persistence of that rate across a quarter century reflects an unresolved accountability structure rather than a clinician compliance problem or a calibration failure. The alerting architecture established in the foundational order entry literature fires a rule against an order and records whether the clinician proceeded. It does not record who authorized the rule, on what basis, or what the clinician determined. The remedy the evidence points toward, patient-specific suppression, creates a new accountability question rather than answering the old one: a system that decides which alerts a clinician never sees is exercising judgment that requires a named owner.

The record

The most rigorously pooled estimate of how often clinicians act on the alerts their systems generate comes from a systematic review and meta-analysis conducted at the Federal University of Santa Catarina in Brazil and published in the Health Informatics Journal in June 2024. Screening 1,873 records across MEDLINE, EMBASE, Web of Science, Scopus, LILACS, and Google Scholar, the reviewers included 16 studies for qualitative analysis and 15 for quantitative analysis. Across 21,435,597 prescriptions in ten studies, the pooled prevalence of a system generating a drug-drug interaction alert was 13.7%, with a 95% confidence interval of 5.6% to 24.7%. Across 570,776 prescriptions in eleven studies, the pooled prevalence of physician override was 90%, with a 95% confidence interval of 85.6% to 95.0%.

The figure is not new and it is not moving.

A systematic review published in the Journal of the American Medical Informatics Association in 2006 by a team at Erasmus University Medical Center in Rotterdam examined 17 papers and found drug safety alerts overridden in 49% to 96% of cases. Eighteen years and an entire generation of systems separate that finding from the 2024 pooled estimate, and the range has not narrowed.

The Rotterdam reviewers interpreted their results through Reason’s framework of accident causation, and located the failure not in the clinician but in what they called error-producing conditions in the alerting system itself: low specificity, low sensitivity, unclear information content, unnecessary workflow disruption, and unsafe handling. That diagnosis is now twenty years old and the conditions it names are still in production.

Neither has tuning helped. Two studies inside the 2024 review measured override before and after the alerting rules were adjusted to reduce volume and raise clinical significance. One moved from 97% to 96%. The other moved from 94% to 74%, an improvement that still leaves three of every four alerts unacted upon. A 2014 study in the same review, examining a commercial interaction system after intensive optimization, concluded that the results suggest the need to question the premises of drug interaction alerting itself.

The explanation most often offered does not survive examination. A 2017 retrospective cohort study at Weill Cornell Medical College analyzed alert data from 112 ambulatory primary care clinicians between January 2010 and June 2013, testing two competing mechanisms: cognitive overload from workload and complexity, and desensitization from repeated exposure. The study found strong effects from work complexity and from repeated alerts for the same patient, and found no evidence of desensitization over time. Acceptance tracked the informational quality of the alert stream rather than how much work the clinician was carrying. In that cohort, acceptance of drug-drug and drug-allergy alerts ran below 1%.

That distinction matters because it changes what kind of problem this is. If clinicians were fatigued, the remedy would be workload. If clinicians were desensitized, the remedy would be novelty, and the literature contains proposals to that effect: retire old alerts, change their presentation, restore salience. The Weill Cornell data support neither. What predicted rejection was repetition of the same alert for the same patient, which is not fatigue but accurate assessment. A clinician who has already determined that an interaction is acceptable for a patient and is told again is not being worn down. They are being told something they have already decided.

The 90% rate is therefore the standing baseline. Any program reporting materially better should be understood as a departure from two decades of pooled evidence rather than as ordinary performance.

The architecture that produces it

The alerting model responsible for that baseline has a traceable origin and a documented success.

A trial at Brigham and Women’s Hospital, published in JAMA in October 1998, established computerized physician order entry as a patient safety standard. Non-intercepted serious medication errors fell 55%, from 10.7 to 4.86 events per 1000 patient-days. The architecture was straightforward: a knowledge base of rules is matched against incoming orders, and an order meeting a rule’s threshold fires an alert.

Two details of that trial are usually left out. The decline was concentrated in potential adverse drug events, which fell 84%. Preventable adverse drug events, the errors that actually reached patients, fell 17%, a result that did not reach statistical significance. And the study tested a second intervention, a clinical team added alongside order entry. The team conferred no additional benefit.

That second finding is worth pausing on, because it was read at the time as evidence that the technology was sufficient. It can equally be read as evidence that the team was given nothing to do. A group convened alongside an alerting system, with no defined authority over what the system flags and no record of what it determined, will not improve outcomes, and the 1998 data say it did not. The conclusion drawn was that the software carried the benefit. The conclusion available is that adding people to a system without giving them a defined role changes nothing, which is a finding about accountability structure rather than about software.

The architecture worked on the measure it was built for. What it does not do, by design, is differentiate among patients. A rule fires identically whether the interacting pair involves a frail patient with impaired renal clearance or a young patient for whom the interaction carries negligible weight, because the architecture matches orders against rules rather than orders against patients.

This is not a defect. It is a consequence of what the system knows. The rule engine holds a drug list and a threshold table. It does not hold the patient’s renal function, the reason the combination was chosen, the monitoring already in place, or the conversation that preceded the order. The clinician holds all of those. The alert therefore arrives carrying strictly less information than the person receiving it, which is an unusual configuration for a safety control, and it explains the override rate more economically than any account of clinician behaviour does. A control that knows less than its operator will be overridden by a competent operator most of the time. That is not a failure of the operator.

That is a description of production systems today. A bibliometric review published in Healthcare in March 2022 by a team at Taipei Medical University, examining 728 articles published between 2011 and 2021 with 24 selected for content analysis, found the research direction shifting from patient safety toward system utility, and identified context-aware design as the field’s next stage. The 2024 meta-analysis reaches the same conclusion from the other direction: adjustment within the existing paradigm produced marginal movement, which is evidence that the limiting factor is the architecture rather than its calibration.

Twenty-six years separate the Brigham trial from that meta-analysis. In that period the underlying design has not changed, the override rate has not changed, and the recommended correction has been stated repeatedly without being built.

What the record does not establish

Three qualifications sit inside these same sources, and a paper arguing from them cannot leave them out.

Overrides are frequently appropriate. The 2006 Rotterdam review states plainly that overriding may often be justified and that adverse drug events arising from overridden alerts are not always preventable. Its authors distinguish between an alert that is appropriate and an alert that is useful, which is a distinction the raw override rate cannot make.

The measured harm is low in monitored settings. Studies inside the 2024 review that followed what happened after an override found adverse drug events in 4% to 7.8% of cases, and in one study none at all. The reviewers offer the obvious reading: in a hospital where patients are monitored continuously, a physician may reasonably judge the benefit to exceed a managed risk. Their argument for concern rests on extrapolation to settings without that support, not on observed harm in the settings studied.

The alerts are a small fraction of prescribing. At a 13.7% generation rate, most orders never produce an alert at all. The override figure describes behaviour at a narrow junction, not a general disposition toward safety information.

None of this rescues the architecture. The 2024 review cites work by Wickens and Dixon establishing that below roughly 70% diagnostic reliability, automation performs worse than no automation at all. A system operating at 10% acceptance is not a safety control that clinicians are misusing. It is a control that has stopped functioning as one, whatever the merits of any individual override.

It also does not rescue the record. An override that was clinically correct and an override that was careless produce identical entries. That is the point of this paper, and the qualifications above sharpen it rather than soften it: if most overrides are justified, then the institution holding that log cannot demonstrate which ones were.

Where the gap actually sits

The override rate is a measurement of a control firing. It is not a measurement of whether anyone was accountable for the response.

Consider what the record contains after an alert is overridden. It contains the order, the alert, and the override. Sometimes it contains a reason code, and the 2024 review notes that in one study the most frequently selected reason was “Other,” with no explanation entered. What the record does not contain is who authorized that rule to fire in that context, on what clinical basis, or what the clinician determined about this patient that the rule could not see.

This is The Accountability Gap™ (TAG™) at the alerting layer, and it has been there since 1998.

The framework defines two seats. A Governance Owner carries institutional accountability for putting a system into use, holding Charter, Commission, and Cover. A Decision Owner carries accountability for the individual call the system informs, holding Decide, Document, and Defend. The gap closes only when both seats are named and The Handoff between them is documented.

The 1998 architecture fills neither seat and documents no handoff. No name attaches to the decision that this rule fires at this threshold in this unit. No record captures what the clinician determined beyond the fact that they proceeded. The override is logged; the judgment is not. An examiner reconstructing the decision a year later finds an event, and no owner on either side of it.

The Governance Owner seat is the one institutions believe they have filled. Most health systems running order entry have a committee that approves the interaction rule set, reviews alert volume, and signs off on changes. That committee holds something close to Charter. It does not hold Cover, because Cover is answering for the chartered operation of the system under examination, and a committee cannot be examined. It has no memory that survives its membership, no capacity to explain why a threshold was set where it was three years ago, and no standing to suspend anything between meetings.

The Decision Owner seat looks better filled and is not. The clinician’s name is on the order, which is why institutions assume accountability is settled. But Decide, Document, and Defend are three functions, and the architecture supports only the first. The clinician decides. The system documents that they proceeded, which is not the same as documenting what they decided. And when the decision is examined, the clinician defends a determination for which the record holds no reasoning, against an alert whose basis they were never shown.

The Named Owner Principle holds that every deployment requires two named persons in the audit trail, that neither substitutes for the other, and that the requirement is satisfied only when both names are recorded. Twenty-five years of alerting has satisfied none of it.

What suppression requires

The correction the evidence points toward is patient-specific alerting. The 2024 review cites a 2022 Saudi study estimating that integrating five laboratory parameters, being potassium, white blood cell count, international normalized ratio, therapeutic drug monitoring levels, and glomerular filtration rate, could eliminate roughly 30% of alerts not clinically actionable for the patient receiving them. The 2006 Rotterdam review had identified substantially the same requirement, naming age, sex, body weight, documented allergies, mitigating circumstances, and drug serum levels. The direction has been stable for eighteen years.

It is also, on its own, insufficient, and for a reason that is easy to miss.

A rule that fires uniformly carries no discretion to misplace. A system that decides, patient by patient, which alerts to surface and which to suppress is making a clinical determination before any clinician sees anything. It is doing so at scale, continuously, and invisibly. The alert that was suppressed leaves no trace unless the system was built to record the suppression itself.

Context-aware suppression therefore converts an accountability gap into a larger one. The 1998 architecture at least produced an artifact: an alert, an override, a timestamp. A suppression architecture produces silence, and silence cannot be audited.

The asymmetry is worth stating precisely. Under the current design, a clinician who overrides an alert has at minimum seen it, and the institution can establish that the information reached the person who acted. Under suppression, the institution can establish nothing of the kind. Asked whether a clinician was informed of an interaction, it can say only that the system determined they did not need to be. That determination was made by logic somebody wrote, on parameters somebody selected, against thresholds somebody set, and unless those decisions carry names the institution is describing an event with no author.

The failure mode is also quieter. An alerting system that is too noisy announces itself: the override rate rises, clinicians complain, the volume is visible in every log. A suppression system that is miscalibrated announces nothing. Alerts that should have fired simply do not, and the absence looks identical to correct operation. Drift in a suppression model is discovered when a patient is harmed, not when a dashboard turns red, because there is no dashboard for events that did not occur.

What the Governance Owner must charter is the suppression boundary itself. Not the alert thresholds, which is what governance committees currently approve, but the specification of what patient-specific logic the system may apply on its own authority and what must remain a clinician’s determination. That specification is written before the model is built, and it carries a name. A boundary defined after deployment is a description of what was built rather than an authorization of it.

What the Decision Owner must hold is the deployed instance. A single enterprise approval covering an alerting portfolio does not create a Decision Owner for each setting the logic operates in. Each deployment requires a named clinician accountable for reviewing what the system suppressed, monitoring calibration as the patient population shifts, and suspending the logic when it drifts. The review of suppressed alerts is the substantive duty here, and it has no counterpart in current practice, because under the present architecture there is nothing suppressed to review.

What The Handoff must produce is a record at the clinical moment. Which patient-specific parameters modulated the alert, what the system determined, and what the clinician did in response. Under the 1998 architecture this record was merely absent. Under suppression it is the only evidence that a decision occurred at all.

A committee is not a name. Health system AI governance is overwhelmingly committee-structured, and a committee can charter a boundary, but it cannot be deposed, cannot be asked what it determined about a particular patient, and cannot suspend a deployment on a Tuesday afternoon.

None of this requires new infrastructure. The parameters are in the record, the logic is buildable, and the logging is a field. What it requires is that the institution decide, before it deploys, who answers for what the system determines on its own.

What this paper is not

Not a claim that overrides are clinically wrong. The evidence indicates many are justified and that measured harm in monitored settings is low. The argument concerns what the record shows, not what the clinician decided.

Not an argument for fewer alerts. Volume reduction has been attempted for two decades and has moved override rates marginally. The argument is that calibration and accountability are different properties and that solving one does not address the other.

Not a claim that context-aware alerting is undesirable. It is the correction the evidence supports. It also introduces a governance requirement the current architecture never had to meet.

Not satisfied by vendor accountability. A vendor supplying suppression logic does not hold Charter, Commission, or Cover for a deployment inside an institution, and cannot.

Not a technical specification. What suppression logic should encode is a clinical and institutional determination. This paper argues that the determination requires a named owner, not what the determination should be.

Not generalizable beyond drug-drug interaction alerting without qualification. Override behaviour for allergy, dosing, and duplicate-order alerts has been studied separately and may differ.

Frequently asked questions

Does a 90% override rate mean the alerts are wrong?
No. It means the control is not producing an owned clinical response. The 2006 Rotterdam review found overriding often justified, and studies measuring what followed an override found low rates of adverse events in monitored hospital settings. The failure is in what the record shows about accountability, not in the clinician's judgment.
If measured harm is low, why does this matter?
Because low measured harm in a continuously monitored hospital is not evidence of low risk in settings without that monitoring, which is the reviewers' own reading. And because a control that cannot be audited will not withstand examination regardless of whether harm occurred.
Has anyone succeeded in lowering override rates?
Marginally, and by mechanisms that carry their own costs. One study in the 2024 review reported lower overrides in a system requiring pharmacist consultation for an override password, which raises the cost of proceeding rather than improving the alert. Two Korean studies reported lower rates in a reimbursement environment penalizing flagged prescriptions.
Is this an argument against clinical decision support?
No. The 1998 trial establishes that order entry reduced non-intercepted serious medication errors by more than half. The argument is that a control which fires and a decision which is owned are separate things, and that the architecture solved only the first.
Who is the Governance Owner for an alerting system in practice?
Whoever the institution names, with authority to charter what the system may determine on its own. In most health systems today that authority sits with a committee, which can hold Charter but cannot hold Cover in the sense the framework requires, because accountability under examination attaches to persons.
What changes if suppression is deployed without these controls?
The institution loses the artifact it currently has. An override is at least a recorded event. A suppressed alert is nothing at all unless the system was built to record it, and no retrospective review can reconstruct a determination that was never written down.
Does this apply to clinical AI beyond alerting?
The mechanism is the same wherever a system produces an output that a named person must act on. This paper treats alerting because alerting has twenty-five years of measurement behind it and therefore carries evidence the newer categories do not yet have.

Funding

None declared.

Conflicts of interest

Mo Johnson, MD MBA is the founder of GPe Research. The Accountability Gap™ (TAG™) and the frameworks named in this paper are works of GPe Research. The author has a commercial interest in the adoption of the frameworks described in this paper.

Published under CC-BY-4.0. Free to share and adapt with attribution.

How to cite

Johnson, M. (2026). Overridden: What a quarter century of ignored alerts establishes about accountability (Position Paper №03). GPe Research Publications. https://publications.gperesearch.com/papers/overridden

Version history

Original publication.

CITE THIS PAPER
APA
Johnson, M. (2026). Overridden: What a quarter century of ignored alerts establishes about accountability (Position Paper №03). GPe Research Publications. https://publications.gperesearch.com/papers/overridden
AMA
Johnson M. Overridden: what a quarter century of ignored alerts establishes about accountability. GPe Research Publications. Position Paper No. 03. Published September 8, 2026. Accessed [date]. https://publications.gperesearch.com/papers/overridden
Chicago
Johnson, Mo. "Overridden: What a Quarter Century of Ignored Alerts Establishes About Accountability." Position Paper №03. GPe Research Publications, September 8, 2026. https://publications.gperesearch.com/papers/overridden.
Vancouver
Johnson M. Overridden: what a quarter century of ignored alerts establishes about accountability [Internet]. GPe Research Publications; 2026 Sep 8 [cited YYYY Mon DD]. (Position Paper; №03). Available from: https://publications.gperesearch.com/papers/overridden
BibTeX
@techreport{johnson2026overridden,
author = {Johnson, Mo},
title = {Overridden: What a Quarter Century of Ignored Alerts Establishes About Accountability},
institution = {GPe Research Publications},
type = {Position Paper},
number = {№03},
year = {2026},
month = {9},
url = {https://publications.gperesearch.com/papers/overridden}
}