The Accountability Canvas™ is a nine block diagnostic and a three layer Gap Score™ for surfacing, scoring, and closing the accountability gap in any regulated AI deployment, before an output reaches the person who has to act on it. It generalizes the instrument first published as Mind the 9 Blocks™ for clinical AI. The nine blocks and the scoring method do not change by sector; only the vocabulary inside each block changes to match the room. This brief states each block in room neutral terms and then shows the clinical and financial wording underneath it, applies the three layer Gap Score™ the same way, and runs two condensed worked examples side by side. It cites Framework Brief №02 for the mechanics of the two named seats and the Handoff, and Working Paper №02 for the financial case it draws on. Mind the 9 Blocks™ remains the canonical clinical language version of the instrument and is not altered by this paper.
The Accountability Canvas™ is a nine block diagnostic and a three layer Gap Score™ for surfacing, scoring, and closing the accountability gap in any regulated AI deployment, before an output reaches the person who has to act on it. It generalizes the instrument first published as Mind the 9 Blocks™ for clinical AI. The nine blocks and the scoring method do not change by sector. Only the vocabulary inside each block changes to match the room.
This brief states each block in room neutral terms first and then shows the clinical and financial wording underneath it, so that an assessor in either room can run the same instrument against a real deployment. The clinical wording is the language already published in Mind the 9 Blocks™. The financial wording follows the case treated in Working Paper №02. The point of setting them side by side is to show that the canvas was always general in structure, and clinical only in vocabulary.
Relationship to Framework Brief №02 and Mind the 9 Blocks™
Framework Brief №02, The Accountability Gap™ states what must be true for a regulated AI decision to hold: two named seats, a Governance Owner and a Decision Owner, each holding three functions, with a documented Handoff between them. That brief is the standard. It does not, on its own, tell an institution whether the standard is met in a given deployment.
The Accountability Canvas is the instrument that answers that question. An institution runs the canvas to find out whether TAG™ is satisfied in a specific deployment, and it runs the Gap Score™ to price how far the deployment stands from satisfying it. The standard names the condition. The canvas measures the distance to it.
Mind the 9 Blocks™ is the first published application of this instrument, written for clinical AI and specific to the bedside. It remains the canonical clinical language version of the canvas, unchanged and separately citable. This paper does not replace it. It states the general instrument that Mind the 9 Blocks™ was the first application of, so that a financial institution, or any regulated institution, has a version it can run in its own room. Where Mind the 9 Blocks™ says patient, bedside, and clinician, the general canvas says the party the decision protects, the point of decision, and the named person who acts.
The instrument matters now because the systems it governs are already acting. Regulated institutions in both rooms are moving AI from pilots into production, and automated systems are disposing of decisions at a volume no person reviews in full. The canvas is built to be run before an output reaches the person who has to answer for it, on the schedule the deployment sets rather than the schedule of the next adverse event or the next examination. An instrument consulted only after an outcome has made the gap visible has become a postmortem rather than a diagnostic.
The nine blocks, generalized
The canvas is read as a whole before it is read block by block. Nine blocks sit in three bands. The first band, Diagnose, asks whether the deployment is even set up to be governed as regulated AI. The second, Assess, asks what is happening at the point of decision and whether the institution can see it. The third, Close, asks who answers for the decision and whether the answer holds over time. The bands are sequential by accountability logic. An institution cannot assess a deployment it has not first distinguished as regulated AI, and it cannot close a gap it has not first assessed.
Each block is stated in room neutral language, then in its clinical and financial applications. A block counts as filled only when a documented, owned, and current artifact stands behind it. A block that is technically present but unowned, or documented but stale, counts as empty.
Diagnose. Assess. Close.
© 2026 GPe Research · CC BY 4.0
Band I: Diagnose
Block 1: AI Distinction. General. Whether the institution governs this deployment as a decision shaping AI system with its own accountability layer, distinct from general enterprise AI. Most institutions apply the general discipline to a system that shapes a regulated decision and never build the distinct layer that decision requires. Clinical AI. Governed as clinical AI protecting the patient decision at the bedside, under a clinical AI policy distinct from the general AI policy, rather than logged in the general AI inventory and passed through an IT review. Financial AI. Governed as financial AI protecting a regulated financial decision, the disposition of an alert, a triage outcome, a fiduciary review, under a policy distinct from the general enterprise AI policy, rather than treated as one more tool in the monitoring stack. What empty looks like. The deployment sits in the general AI inventory, cleared by a technology review, with no one having asked whether a system that shapes a regulated decision needs governance an operational system does not. What filled looks like. The deployment is formally classified as decision shaping AI, governed under a policy distinct from the general AI policy, with the decision it protects named as the thing being governed.
Block 2: Governance Cascade. General. Whether the layered governance the deployment depends on is actually built, ending in the decision layer where the output becomes an institutional act. Each layer depends on the one above it, the layers converge at the decision, and the decision layer is the one most institutions have not built, which leaves the center of the convergence exposed. Clinical AI. Data Governance, AI Governance, Healthcare AI Governance, and Clinical AI Governance, with the fourth layer at the bedside as the one usually unbuilt and the one the named owner sits in. Financial AI. Data Governance, AI Governance, and the institution’s financial and regulatory governance, converging at the point where an automated output becomes a financial decision, with that decision layer, the alert queue or the review desk, as the one usually unbuilt. What empty looks like. A mature data program and a general AI program are treated as coverage for the decision layer, which is named in no policy and operated by no one. What filled looks like. Every layer exists, each documented and each with named accountability, and the dependency between them is explicit, so a failure at the decision layer is not masked by maturity at the data layer.
Block 3: Affected Populations. General. Who the deployment affects, at the level of the institution’s actual population rather than the vendor’s validation cohort, and whether anyone is accountable for detecting performance that diverges across groups or conditions. Clinical AI. The patient population, the demographic gap between the validation data and the patients the model now informs decisions about, with a named owner for disparate impact monitoring. Financial AI. The customer and transaction population, the gap between the volume and risk profile the system was tuned for and the book of business the institution actually carries, with a named owner for monitoring performance across products, segments, and geographies. This is the gap Working Paper №02 documents directly, where thresholds set for a smaller institution were never retuned to the business the bank had grown into. What empty looks like. Aggregate performance was accepted at a pilot, no analysis was run across the groups or conditions the deployment actually spans, and no one is named to monitor for divergence after it goes into use. What filled looks like. A written definition of the affected population per deployment, compared to the data the system was validated on, with a named owner accountable when performance diverges across groups or conditions.
Band II: Assess
Block 4: The Promise. General. What the deployment is expected to deliver, defined by the institution rather than the vendor, with a baseline, a metric, a horizon, and a named sponsor. The vendor describes what a system can do in general. The promise documents what this institution expects it to do specifically. Clinical AI. The clinical promise, what improvement this institution expects in this patient population against this baseline over this horizon, signed by a named clinical sponsor. Financial AI. The institution’s defined expectation for what the system delivers. For an alert triage system, what it should escalate and what it may close at the institution’s actual volume and risk, with a baseline, a metric, a review date, and a named risk or compliance sponsor. What empty looks like. The deployment was approved on vendor material or peer adoption, with no baseline, no defined metric, and no named sponsor who stated what the institution expected it to deliver. What filled looks like. A written statement produced before the deployment goes into use, with a baseline, a defined metric, a horizon, and a named sponsor accountable for whether the system delivers it.
Block 5: The Three Decisions. General. Whether the record can distinguish patterns that look identical after the fact but are not the same decision: a decision made on the person’s judgment alone, a decision an AI output shaped, and a decision where the person and the system reached the same conclusion independently. A fourth pattern, where the system disposed of the matter and no person decided at all, is the one this block is most built to surface. Clinical AI. A Clinician Decision, a Shaped Decision, and a Parallel Decision, distinguishable in the chart without reconstruction. Financial AI. The officer’s decision, the shaped decision, the parallel decision, and the auto disposition where the system closed an alert and no person decided, distinguishable in the alert record. Working Paper №02 documents the failure mode directly, where a triage system closed a very high percentage of ingested alerts and produced a volume of dispositions no person had made. What empty looks like. The output appears in the workflow, the person acts, and the record shows only the action, so whether the system shaped the decision, was overridden, or merely agreed is invisible afterward. What filled looks like. The output is captured at the point of use, the person’s response is recorded as acceptance, override, or independent concurrence, and the patterns are distinguishable in the record without reconstruction.
Block 6: Accountability Trail. General. The governed record that links the AI output to the decision to the outcome, in a form governance can review without a technical reconstruction. A log records what a system did. The trail records what the institution did, who was accountable, what they knew, what they decided, and what followed. Clinical AI. Recommendation to clinical decision to patient outcome, logged in the clinical record at the moment of use and reviewable within a defined window rather than reassembled under subpoena. Financial AI. Alert or output to the officer’s disposition and rationale to the outcome, a filing, an account action, or a documented closure, reviewable by an examiner without reassembling it from separate system logs. What empty looks like. The output is logged in one system and the decision in another, with no link between them and no connection to the outcome, so reconstruction after an event takes weeks and is often incomplete. What filled looks like. The output is logged where the decision is made, the person’s response and reasoning are captured as structured data, the outcome is linked to the decision, and the whole trail is reviewable within a defined window.
Band III: Close
Block 7: The Named Owners. General. The two named seats the Named Owner Principle requires, carried on the deployment’s record, each holding its three functions. The mechanics are set out in Framework Brief №02 and are not restated here. Both names must appear, neither substitutes for the other, and a seat that holds none of its functions is a name on a record, not an owner. Clinical AI. A Governance Owner, the Chief Medical Officer or Chief Medical Information Officer, holding Charter, Commission, and Cover; and a Decision Owner, the treating clinician, holding Decide, Document, and Defend. Financial AI. A Governance Owner, the Chief Risk Officer or Chief Compliance Officer, holding Charter, Commission, and Cover; and a Decision Owner, the officer or fiduciary who reviews the escalated output, holding Decide, Document, and Defend. What empty looks like. One seat is named and the other is left implicit, or the record shows the person who acted but nothing about the authority that approved the system, with the gap between them undocumented. What filled looks like. Both seats named, both recorded on the deployment’s record, both current as of its present scope, and each holding all three of its functions.
Block 8: Governance Actions and Partners. General. Two accountabilities that share a block because they share a failure mode: the formal institutional actions that create a governance record, with dates and signatories, and the external relationships that share accountability before an event rather than during one. Clinical AI. A clinical AI policy adopted by the medical staff, documented review cadence, and board reporting; with vendor contracts, malpractice carrier briefings, regulator engagement, and counsel that allocate accountability before it is tested. Financial AI. Board and committee actions, adopted policy, and reporting on the record; with vendor contracts, regulator engagement, insurer briefings, and counsel. The insurance dimension is treated in Position Paper №02, which documents how the standard forms have begun to exclude losses arising from generative AI. What empty looks like. Governance lives in steering committees and email threads, with no adopted policy and no board visibility, and the vendor, the insurer, and counsel were never engaged to share the accountability before an event. What filled looks like. An adopted policy, a documented review cadence, board reporting on the record, and vendor contracts, insurer briefings, and legal engagement that allocate accountability before it is tested.
Block 9: Risk Exposure & Gap Score™. General. The synthesis. The institution’s aggregate view of its regulated AI risk: every deployment catalogued, each block assessed, and the Gap Score™ that states where the institution stands. This block cannot be filled until the other eight have been assessed with substance, because it is the measure of them. Clinical AI and Financial AI. An enterprise register of deployments, a Gap Score™ calculated per deployment and in aggregate, reviewed on a defined cadence, with a board level summary and a remediation roadmap sequenced by exposure. The register is the same instrument in either room. Only the deployments listed in it differ. What empty looks like. Each deployment is governed in isolation, no enterprise register exists, the Gap Score™ has never been calculated, and regulated AI does not appear as a named category in enterprise risk reporting. What filled looks like. An enterprise register of deployments, a Gap Score™ per deployment and in aggregate, reviewed on a defined cadence, with a board level summary and a remediation roadmap sequenced by exposure.
The Gap Score™ methodology, generalized
The Gap Score™ converts a deployment’s accountability from a matter of judgment into a defensible number, produced before the output reaches the person who acts on it rather than reconstructed after an outcome. Where the nine blocks diagnose what a deployment requires, the Gap Score™ prices whether the deployment is ready to sign off. It is scored across the three layers that must converge for a regulated AI decision to hold. The math and the layer structure are identical to the clinical version. Only the applied language and the worked inputs change by room.
Each layer answers one question, and each is scored from 1 to 5.
AI Governance, the enterprise layer, asks whether the enterprise can defend this. The readiness test is whether a documented enterprise governance trail exists for this specific deployment, not a policy that covers AI in general. This layer is scored on the governance trail, the audit documentation, and the regulatory alignment behind the specific deployment, and it reads the same way in either room.
The domain protection layer asks whether the deployment protects the party the decision exists to protect, validated against the actual population and conditions it operates in. In clinical AI this layer is Clinical AI Governance, the patient safety layer, scored on population validation, edge case testing, and clinical standard alignment. In financial AI this layer is Financial AI Governance, the customer and market integrity layer, scored on whether the system is calibrated to the institution’s actual transaction volume and risk, whether its tuning has been validated against the current book, and whether it covers the higher risk products and cross border exposure the institution has taken on. The question is identical. The population being protected is what changes.
Decision Accountability, the named owner layer, asks whether the owners are named. The readiness test is whether both named owners exist before the first output reaches the person who acts on it, and whether the six functions are held: Charter, Commission, and Cover at the institutional seat, and Decide, Document, and Defend at the seat of the person who acts. Not a committee. Not a vendor contact. This layer scores low when only one seat is filled and fails when neither is, and a seat that is named but holds none of its functions scores as unfilled. It reads the same way in either room.
The three scores sum to a total out of 15, and the total sets the institution’s position. A total above 12 is defensible: the owners are named, the decision is traceable, and the governance trail holds. A total below 9 is unpriced liability: the institution is holding risk it has not measured and cannot defend. The range between 9 and 12 is the contested middle, defensible only with intervention before the deployment goes into use. The discipline is a gate, not a gauge. If the institution cannot complete the score across all three layers before the output reaches the person who acts on it, the deployment is not ready.
The score sits on top of a way of running the canvas, and the way is the same in either room. Three operations move an institution through the nine blocks. The cascade architects the canvas, building the governance layers the Diagnose band assesses, because an institution cannot assess or close what it has not first built. The evaluation produces the Gap Score™, rating the three convergence layers rather than the nine blocks directly, which is why the canvas and the score are two instruments and not one. And a repeatable loop runs it deployment by deployment: categorize the deployment as regulated AI, locate the point where the output meets the decision, own both named seats, score the three layers, and evaluate the trail the deployment produces. An institution does not fill the canvas once. It runs the loop on every deployment and runs it again whenever the deployment changes.
Underneath the operations sits a test that separates a block that is substantively filled from one that is merely documented. A block counts as filled only when the institution can validate that the system performs for the population it actually serves, assign a named owner to the accountability, link the output to the decision and the outcome, uncover the gaps the deployment introduces rather than waiting for them to surface, and embed the governance in the workflow rather than bolting it on. A block whose paperwork exists but whose conditions do not hold is an empty block with a document in front of it. The most common failure of governance is not absence. It is the appearance of presence.
Two worked examples, side by side
The two runs below are condensed and illustrative. They show how an assessor reaches a score in each room, not an official rating of any institution. The clinical run adapts an example of the kind treated in Mind the 9 Blocks™. The financial run reads the pattern documented in Working Paper №02 through the same three layers, and cites that paper for the underlying case rather than restating its evidence.
Clinical run: a sepsis prediction model on a medical service. The model is validated and monitored, and an enterprise AI policy exists, but the governance trail specific to this deployment is thin, so AI Governance scores in the middle. The training population has been compared to the service’s actual patients only in the aggregate, with no subgroup analysis, so the domain protection layer scores in the lower middle. A Governance Owner is named in the charter, but the bedside Decision Owner is left implicit and the six functions are not all held, so Decision Accountability scores low. The total lands below the defensible line, which tells the institution the deployment is not ready to sign off and shows exactly which layer to close first.
Financial run: an automated alert triage system, on the pattern in Working Paper №02. A general enterprise AI policy exists but no governance trail is documented for the triage deployment itself, so AI Governance scores low. The system’s thresholds were never tuned to the institution’s actual transaction volume and risk, so the domain protection layer, Financial AI Governance, scores at the floor. The system auto closed a very high percentage of ingested alerts, so for those dispositions the Decision Owner never received the alert and the seat went unfilled by construction, so Decision Accountability scores at the floor. The total sits deep in unpriced territory, which is a reading consistent with the supervisory outcome.
Neither run is a verdict on the model. The clinical model may be accurate and the financial system may be technically sound, and both still score low, because the score measures whether the institution can name who owns the decision, protect the population the decision serves, and defend the trail, not whether the model performs. The value of a run is that it names the layer to close first. In the clinical run that is Decision Accountability, the implicit bedside seat. In the financial run it is the domain protection layer, the untuned thresholds, followed closely by the reviewing seat the auto close left empty. A score is not a grade to file. It is a sequence of work, ordered by exposure.
| Layer (scored 1 to 5) | Clinical run | Financial run |
|---|---|---|
| AI Governance (enterprise) | Middle: policy exists, deployment specific trail thin | Low: no documented trail for the triage deployment |
| Domain protection | Lower middle: no subgroup validation | Floor: thresholds never tuned to actual volume and risk |
| Decision Accountability | Low: one seat implicit, functions not all held | Floor: auto close leaves the Decision Owner seat unfilled |
| Position | Below the defensible line | Deep in unpriced liability |
The instrument reaches a defensible reading in both rooms using the same three layers and the same bands. What differs is the vocabulary of the inputs, not the method that scores them.
Why the instrument generalizes without modification
The canvas and the scoring method were never clinical in structure. They were clinical only in vocabulary. The nine blocks measure institutional and procedural conditions that exist identically in any regulated AI deployment: whether the deployment is distinguished as decision shaping AI, whether the governance layers beneath the decision are built, whether the affected population is defined and monitored, whether the institution has stated what it expects the system to deliver, whether the record can tell the decisions apart, whether the trail links output to decision to outcome, whether the two seats are named and their functions held, whether formal actions and external partners are secured, and whether the whole is catalogued and scored.
None of those conditions is medical. Each is a property of how an institution governs a decision that an AI output shapes and a named person has to answer for. The bedside is one instance of the point of decision. The alert queue is another. The instrument reads them with the same nine blocks because the accountability structure underneath them is the same structure.
This portability is a structural property being described, not a business claim being made. The two rooms with published applications are clinical AI and financial AI. Where any other regulated sector runs an AI system that produces an output a named person must act on, the same nine blocks and the same three layers apply, and the institution that has run the canvas in one room already holds the instrument it needs in the next.
The proof of that claim is the third section of this brief. Nine blocks were stated in room neutral language and then filled with clinical and financial wording, and not one block had to be added, removed, or restructured to hold both rooms. The vocabulary changed in every block. The block changed in none of them. An instrument that survives that substitution intact was measuring something underneath the vocabulary all along, and what it measures is the accountability structure a regulated institution owes for a decision an AI output shapes and a named person has to answer for.
Framework glossary
The Accountability Canvas. A nine block diagnostic and a three layer Gap Score™ for surfacing, scoring, and closing the accountability gap in any regulated AI deployment, generalized from its first clinical application, Mind the 9 Blocks™.
Gap Score™. The three layer score, from 1 to 5 on each of AI Governance, the domain protection layer, and Decision Accountability, for a total out of 15, that prices how far a deployment stands from a defensible accountability position before the output reaches the person who acts on it.
Block 1, AI Distinction. Whether the deployment is governed as decision shaping AI with its own accountability layer, distinct from general enterprise AI.
Block 2, Governance Cascade. Whether the layered governance the deployment depends on is built through to the decision layer where the output becomes an institutional act.
Block 3, Affected Populations. Who the deployment affects, and whether performance across groups and conditions is defined and monitored against the institution’s actual population.
Block 4, The Promise. What the institution, not the vendor, expects the deployment to deliver, with a baseline, a metric, a horizon, and a named sponsor.
Block 5, The Three Decisions. Whether the record can distinguish a human decision, a shaped decision, a parallel decision, and a disposition where the system acted and no person decided.
Block 6, Accountability Trail. The governed record linking the AI output to the decision to the outcome, reviewable without a technical reconstruction.
Block 7, The Named Owners. The two named seats and their six functions that the Named Owner Principle and The Accountability Gap™ require.
Block 8, Governance Actions and Partners. The formal governance actions on the record and the external relationships that share accountability before an event.
Block 9, Risk Exposure & Gap Score™. The aggregate register and score that state where the institution stands across all of its regulated AI deployments.
Frequently asked questions
- What is the difference between the Accountability Canvas and Mind the 9 Blocks™?
- They are the same instrument at two levels. Mind the 9 Blocks™ is the first published application, written entirely in clinical language for clinical AI, and it remains the canonical clinical version. The Accountability Canvas is the room neutral version underneath it: the same nine blocks and the same three layer Gap Score™, stated so that any regulated sector can run them. Mind the 9 Blocks™ is not changed or replaced by this paper.
- Can we run the Gap Score™ on a financial AI deployment today?
- Yes. The three layers and the scoring math are identical to the clinical version. Section 4 states each layer in room neutral terms and shows how a financial assessor applies it, and section 5 runs a condensed financial example. An institution scores AI Governance, the domain protection layer, and Decision Accountability from 1 to 5 each, for a total out of 15, before the output reaches the person who acts on it.
- Does a low Gap Score™ mean the AI system itself is unsafe?
- No. The Gap Score™ measures accountability, not model performance. A low score means the institution cannot yet name who owns the decision, protect the population the decision serves, or defend the governance trail, regardless of how accurate the model is. A capable model with a low Gap Score™ is a well built system nobody is accountable for.
- Who inside an institution should run this assessment?
- The person who will have to answer for the deployment, working with the seats the canvas names. In a health system that is typically the Chief Medical Officer or Chief Medical Information Officer with the General Counsel and Chief Risk Officer. In a financial institution it is typically the Chief Risk Officer or Chief Compliance Officer with counsel. The score is reported to the board.
- How does this relate to TAG™ and the two named seats?
- The Accountability Gap™ names what must be true: two named seats, six functions, and a documented Handoff between them. The Accountability Canvas is the instrument an institution runs to find out whether that is true in a specific deployment, and the Gap Score™ prices how far it stands from being true. TAG™ is the standard. The canvas is the measurement.
- Is this canvas a compliance requirement or a diagnostic tool?
- It is a diagnostic tool. It describes the accountability structure that regulators, boards, and courts already expect an institution to be able to show, and it measures the distance to it. Running it is voluntary. Being asked who owned a decision is not.
- How often should an institution re run the Gap Score™ on a given deployment?
- Whenever the deployment changes and on a fixed cadence in between. Owners rotate, volumes and populations shift, and thresholds set once decay against a moving baseline. A score that was defensible at go live describes accountability that may no longer exist, so the canvas is run again at each material change and reviewed on a schedule the deployment sets.
Funding
None declared.
Conflicts of interest
Mo Johnson, MD MBA is the founder of GPe Research. The Accountability Gap™ (TAG™), the Clinical AI Accountability Canvas™, Mind the 9 Blocks™, the Gap Score™, the Named Owner Principle, MedicoVigilance™, and FinVigilance™ are works and marks of GPe Research. The author has a commercial interest in the adoption of the frameworks described in this brief.
How to cite
@techreport{johnson2026accountabilitycanvas,
author = {Johnson, Mo},
title = {The Accountability Canvas: A Nine Block Diagnostic for Any Regulated AI Deployment},
institution = {GPe Research Publications},
type = {Framework Brief},
number = {№03},
year = {2026},
month = {7},
url = {https://publications.gperesearch.com/papers/the-accountability-canvas}
}