Health Reclaimed · Reader resources

Appendices A–D

The four assessments from the back of the book, in full. Each one asks what an institution would have to be able to do before an artificial intelligence system could produce the outcome it was bought for.

  • Free to read
  • No sign-in
  • 2,160 words

Appendix A · Readiness Assessment

What an organization should be able to answer before it acquires an AI system.

Before acquiring an artificial intelligence system, an organization should be able to answer a narrower question than whether it is ready. It should be able to answer what it would do if the system worked.

This assessment is built around the five conditions described in the introduction. A system that performs well can still fail if any one of them is missing.

Work through each section and record what you actually know, not what you assume. Where you cannot answer, write that you cannot. An unanswered question is the most useful result this assessment produces.

The five conditions in series A working system passes through legal standing, financial capacity, procedural clarity, contractual position and legitimacy before an outcome reaches anyone. Financial capacity is shown absent, and the path ends at that point. System performs correctly Legal standing Financial capacity absent path ends here Procedural clarity Contractual position Legitimacy Outcome reaches people
The five conditions run in series, not in parallel. A system can perform perfectly and still produce nothing if one condition is missing — here, the capacity to act on what it finds. Downstream conditions may be fully in place and never get reached.

Legal standing

  • Does your organization have the authority to act on what this system would produce? Name the source of that authority.
  • If the system identifies a risk in a population you do not directly serve, what standing do you have to intervene?
  • If someone challenges an action taken on the basis of this system's output, who defends it and on what basis?
  • Are there actions the system might recommend that your organization is not permitted to take?

Financial capacity

  • What is the cost of acting on the output at expected volume, separate from the cost of the system itself?
  • If the system identifies more cases, more risks, or more need than you currently address, is there capacity to respond, or will the additional findings go unaddressed?
  • Is monitoring funded as an ongoing operating cost rather than a project expense?
  • What happens to the system when the funding that acquired it ends?

Procedural clarity

  • Who receives the output? Name a role, not a department.
  • What is that person required to do with it, and within what period?
  • What happens if they do nothing? Is that visible to anyone?
  • How does this output enter existing workflows, and what does it displace? If it is added to a queue that is already full, what falls off?
  • Who is authorized to override the system, and what record does an override leave?

Contractual position

  • What does your agreement with the supplier require them to disclose about how the system was developed and on what data?
  • What evidence of the system's operation will exist, who holds it, and can you obtain it without the supplier's cooperation?
  • If the system's performance degrades, what obligation does the supplier have, and what remedy do you have?
  • Does your agreement require the supplier to obtain and preserve the information you may need from parties upstream of them?
  • Who negotiated these terms, and what technical documentation did they review before accepting them?

Legitimacy

  • Do the people this system will make determinations about know it exists?
  • Was the community consulted before deployment or informed after it?
  • If someone believes the system was wrong about them, how do they say so, and who is required to respond?
  • Would you be comfortable explaining this deployment publicly, in detail, to the population it affects?

Recording what you found

For each condition, write the questions you could not answer. Do not summarize. Do not score. List them.

Legal standing

  • Financial capacity
  • Procedural clarity
  • Contractual position
  • Legitimacy
  • Using this assessment

Any question you cannot answer is a gap, and gaps in the same section compound. A section where most questions are unanswered indicates a condition that is not in place, which means the system will not produce the outcome you are buying it for regardless of how well it performs.

One revealing pattern is strong answers under legal standing and procedural clarity alongside weak or unanswered questions under financial capacity and legitimacy. That combination describes an organization that can lawfully act, knows who would act, and may have neither the resources to act at scale nor the standing in the community to be trusted when it does.

This assessment is deliberately something you complete about your own organization. That has a limit worth naming. Chapter 12 argues that an evaluation performed by a party with an interest in the result is worth less than one performed independently. The same applies here. Use this to find your gaps, not to certify that you have none.

Appendix B · Equity in Detection and Response

Whether your institution can respond equitably to what an accurate system finds.

Most equity assessments for artificial intelligence ask whether a system is fair. That question matters, and Chapter 10 examines it at length.

This assessment asks a different one. It asks whether your institution can respond equitably to what an accurate system finds.

The distinction is the argument of this book applied to equity. A system that correctly identifies elevated risk in an underserved community, deployed by an institution with no capacity to serve that community, does not close a disparity. It documents one. The detection was accurate. The gap widened anyway.

So the questions below move outward, from the data to the response.

Where equity assessments stop, and where the gap opens Five questions in order. The first two, about representation in the data and performance by group, are where most assessments end. The last three, about response capacity, who bears each kind of error, and the effect on the disparity, are where the gap actually widens. WHERE MOST EQUITY ASSESSMENTS STOP 1  Who the system sees 2  Who it performs for WHERE THE DISPARITY ACTUALLY WIDENS 3  Who you can actually reach 4  Who bears the error 5  What the response does to the gap
The questions move outward, from the data to the response. An assessment that ends after question two can certify a fair system operating inside an institution that cannot serve the population it just identified.

Who the system sees

  • Which populations are represented in the data this system was built on, and which are absent or thinly represented?
  • Were any groups excluded by the way the data was collected rather than by decision? Who is missing because they do not appear in the systems that generated the training data?

Who the system performs for

  • Does performance vary across the demographic and geographic groups you serve? How do you know?
  • If you cannot measure performance separately by group, what would it take to be able to?

Who you can actually reach

This is the section most equity assessments omit, and it is where detection turns into prevention or fails to.

  • For each population the system may identify, what is your capacity to respond? Name the service, the pathway, and the constraint.
  • Is your response capacity distributed the same way your detection capacity is? If the system finds more need in a community you serve least well, what happens?
  • Which findings would you be unable to act on, and for whom?

Who bears the error

Every system produces two kinds of error, and they rarely fall on the same people.

  • Who is harmed when the system flags someone incorrectly? What does that cost them?
  • Who is harmed when it misses someone? What does that cost them?
  • Are those two burdens carried by the same populations, or by different ones?

What the response does to the gap

  • If this system operates as intended for one year, does the disparity you care about narrow or widen? Say why.
  • What would you need to measure to know the answer rather than assume it?

Using this assessment

An accurate system deployed into an institution with uneven response capacity will produce uneven outcomes. That is not a model failure and it will not be caught by testing the model.

The pattern to watch for is strong answers in the first two sections and weak ones in the third. That describes an organization that has thought carefully about whether its system is fair and has not asked whether it can act on what the system finds. It is the most common way an equitable intention produces an inequitable result.

Equity is not established at acquisition. It is a property of what happens after detection, measured over time. Appendix D is where that measurement lives.

Appendix C · Authority Map

Who, specifically, holds the authority to act — and under what instrument.

Appendix A asks whether your organization has the authority to act. This one asks who, specifically, holds it.

Complete one map for each artificial intelligence system in operation. The purpose is not to describe roles in general. It is to be able to name a person for every action the system might require, before the system requires it.

For each action, record four things: who is authorized, under what instrument that authority exists, what record the action leaves, and what happens if nobody acts.

An instrument is a statute, regulation, organizational policy, contract term, or written delegation. "Everyone understands that they would" is not an instrument.

The authority map: eight actions, four facts each A grid of eight actions against four columns. Two rows, suspend or stop the system and obtain the operating record, are highlighted as the ones most often left empty. Who Instrument Record If none Review the output Act on the output Escalate Override or disregard Suspend or stop Obtain the record Change thresholds Speak publicly Highlighted rows are the ones most often left blank.
Each action needs a named person, a named instrument, a record, and a consequence for inaction. A blank instrument column means an authority that exists in practice but nowhere in writing; a blank row means nobody holds that power at all.

The actions to map

For each action below, record four things: who is authorized, under what instrument that authority exists, what record the action leaves, and what happens if nobody acts.

  1. Review the output
  2. Act on the output
  3. Escalate beyond the first responder
  4. Override or disregard the output
  5. Suspend or stop the system
  6. Obtain the operating record
  7. Change the system's configuration or thresholds
  8. Speak publicly about the system's operation

Reading the map

Three patterns are worth noticing when you finish.

Blank instrument fields. An authority that exists in practice but not in any instrument will hold until it is contested. It is the first thing to fail under pressure, and pressure is exactly when it is needed.

One name in every row. A single person authorized to review, act, escalate, override, and stop is not oversight. It is a single point of failure with a title.

Empty rows. The rows most often left blank are suspend or stop, and obtain the operating record. Those are the two authorities that matter most when something has gone wrong, and they are the two least likely to have been established before it did.

Complete this map before deployment. Completing it afterward is possible, but the answers tend to be shaped by what already happened.

Appendix D · Monitoring Framework

What has to be watched after deployment, and who has to be told.

Artificial intelligence systems are not finished when they are deployed. Data shifts. Populations change. Models trained on one period continue operating into another. A system that performed correctly at acquisition can degrade without anyone noticing, because nothing about the degradation announces itself.

Monitoring is how an institution keeps the ability to answer for what a system did. It is not a technical activity. It is the practice that makes accountability possible after the fact.

Four things need to be watched.

Monitoring as a loop that produces a record Performance by group, action taken, evidence available and standing with the affected population feed a fixed-interval review, which produces written findings for whoever is accountable, which are kept as the record. An arrow returns to the top for the next interval. WATCHED CONTINUOUSLY Performance across populations Action taken Evidence available Standing with the affected population at a fixed interval Written findings to whoever is accountable, whether or not anything changed Kept as the record next interval
The loop, not the report, is the product. Findings go out on a fixed schedule whether or not anything changed — a process that reports only exceptions leaves no evidence that monitoring happened at all.

Performance across populations

Not aggregate accuracy. Aggregate accuracy conceals the failure that matters, which is a system that works well overall and poorly for a specific group.

Track results separately by the demographic and geographic categories relevant to your population. Watch for divergence over time, not only at deployment. A system that performed evenly at acquisition can drift unevenly.

Record what you find, including when you find nothing.

  • Are outcomes diverging for any group?
  • Has performance changed since the last review?
  • Do we hold a record of results by group and over time?

Action taken

The measure most often omitted, and the one closest to the argument of this book.

Track what proportion of the system's outputs resulted in an action. Not how many alerts fired. How many produced a decision, a referral, an inspection, a contact.

A system generating findings that nobody acts on is not a functioning system. It is a growing record of institutional knowledge without institutional response, and that record exists whether or not anyone acted.

If the action rate is low, determine whether the cause is output quality, workflow capacity, or absent authority. Those require different remedies.

  • What proportion of outputs led to an action?
  • What kinds of action were taken, and what kinds were not?
  • If the rate is low, is the cause output quality, workflow capacity, or absent authority?

Evidence available

Confirm, at intervals, that you can actually produce the record you would need if the system's operation were questioned.

This check can fail quietly, because the absence of a usable record may not become apparent until someone needs it.

  • Can you retrieve what the system determined, when, and on what basis, without the supplier's assistance?
  • Can that record be verified by someone other than the party that produced it?
  • Does it trace past your immediate supplier to the parties upstream?

Standing with the affected population

Whether the people subject to the system's determinations know it operates, and what they say when asked.

Track appeals, complaints, and refusals. A falling participation rate is a monitoring finding, not a communications problem.

An institution that has lost standing with a community cannot act effectively on what it knows about that community, however accurate its systems are.

  • Do affected people know the system operates?
  • What are appeals and complaints saying?
  • Is participation stable, rising, or falling?

Cadence and destination

Review at a fixed interval that does not depend on anyone remembering. The interval should reflect the system's risk, rate of change, operating environment, and the consequences of failure.

Findings go to whoever is accountable for the system's operation, in writing, whether or not anything has changed. A monitoring process that reports only exceptions produces no evidence that monitoring occurred.

Keep the findings. The purpose of monitoring is not primarily to catch problems, though it does. It is to ensure that when someone asks what you knew and when you knew it, there is an answer.

Review record

  • System reviewed
  • Review period
  • Conducted by
  • Reported to
  • Findings
  • Changes made, or reason none were made

Working with these

Fillable versions, and the book behind them.

The assessments above are the reading versions and stay open. The worksheets you can complete, save and circulate inside your own organization are in the reader community, free with membership.

Free for readers · Hosted on Skool