Artificial Intelligence|The error dividend

Machine, heal thyself

Two 2026 trials show that medical AI can improve the record and amplify bad advice. The contest is over who controls the lessons.

Corrected July 18th: A categorical prediction about clinicians withholding candour was narrowed to the likely effect on the quality of the record.

Editorial illustration: a clinician controls a medical printing press that reproduces both red repair stitches and black ink stains
Illustration generated with OpenAI; art direction by Threadonomist

On June 26th a trial in 16 Kenyan primary-care clinics gave medical AI an unusually practical exam. A large-language-model assistant, built into the electronic record, made clinical notes more comprehensive and more likely to contain an appropriate diagnosis and treatment plan. It did not significantly reduce treatment failure within 14 days.1

A second 2026 trial supplied the warning label. Forty-four doctors in Pakistan, all trained to use AI, diagnosed six simulated cases with optional access to ChatGPT. Half were offered unmodified advice; half saw clinically significant errors in three of the six cases. The second group scored 73.3% for diagnostic reasoning, against 84.9% in the first.2 The trials tested different things. Together they show that software may improve a medical record more readily than a medical outcome—and can carry a mistake as efficiently as a lesson.

That mixed verdict may still change medicine before AI cures anyone. The larger promise is institutional: an error could become cheaper to find, compare with other cases and turn into an alert, a protocol or a better model. Call it the error dividend. A model does not collect it by itself. Its value depends on who may inspect the records, test a change and carry a proven improvement into other clinics. Who owns the AI will matter less than who controls what medicine learns from it.

Aviation built an institution around exactly that idea. For half a century NASA's confidential reporting system has offered pilots a bargain: a qualifying report may spare them a civil penalty or licence suspension. Accidents, crimes and deliberate violations are excluded. The point is candour, early enough that a near-miss stays one.3 Medicine has national incident systems of its own. But reporting alone does not connect an incident with the clinical record, the patient's outcome and a corrective action across the institutions that deliver care. A buried error still harms one patient and teaches few others.

One mistake can teach the system, police the clinician—or disappear into a vendor's model.

The useful mistake

An error must first leave a usable trace. A complaint, review or compensation claim makes an injury visible; an event that leaves none may vanish with the chart. America's National Practitioner Data Bank records malpractice payments. New Zealand and Sweden use no-fault schemes that can compensate injury without first proving a clinician negligent.4 Such schemes are less a price on human harm than a way to produce evidence, and producing it is only the start. Records that cannot be inspected are merely stored; patterns that cannot change practice are merely described. Here “ownership” extends beyond title to a patient's chart. It is a bundle of rights: to inspect incidents, audit the software, improve a model, reuse derived knowledge, move it elsewhere and share in any return.

From one mistake to a shared lesson. Four institutional steps, each requiring a permission the one before it does not. This is a map of permissions, not a claim that a model learns online from one patient.

How those rights are distributed determines what the record becomes. A learning system turns mistakes into safer practice. A control system turns them into rules for auditing clinicians. A fragmented system leaves each hospital and vendor to learn alone. The three will overlap, but they are not equivalent. The same record that warns one doctor can be used to overrule another.

Three fates for the error dividend. Each path carries its own logic, examples and the right that matters most. These are competing tendencies, not forecasts.

Who collects the dividend?

Whether a record warns a doctor or overrules one depends on who holds which right. The Kenyan trial shows how that bundle can be split. Kenyan clinicians and researchers helped design and run the study. Its prompt is open, and de-identified inputs, model outputs and clinical outcomes are due to enter a public repository. Yet the underlying GPT-4o model and the electronic-record system are proprietary, and commercial use of the full product must be negotiated separately.5 Participants can inspect some of the evidence, but they cannot simply carry the full product elsewhere. Such a split is ordinary. What it settles is who may build on the result.

The split matters most where adoption is quick and bargaining power uneven. Rwanda's three-year memorandum with Anthropic covers health, education and public services, but the announcements do not spell out the rights to data or knowledge derived from them. The African Union calls for data sovereignty and locally retained value; the World Health Organization urges governments to seek rights in products built through public-private health projects.6 Without explicit rights, a country could gain a useful tool and local skill while a foreign firm kept the product and the right to sell what it had learned. The real choice is between arrangements that build a learning capacity and those that merely rent one.

China offers the opposite concentration of power. Its 2025 plan calls for broad use of imaging and clinical-decision tools in larger hospitals, backed by stronger health-data infrastructure by 2030.7 Pooling records at that scale could support comparison across every hospital in the country. It would also give the state great influence over who contributes, who benefits and what counts as an error. The plan describes the machinery, not the politics.

Blame, incorporated

The politics begin when the record is used to judge a worker. Retrospective software can flag decisions that differ from a payer's preferred standard. In a teacher's hands that may improve care; in an insurer's it can make a refusal cheaper. An ongoing federal case alleges that UnitedHealth used nH Predict estimates to press for shorter post-acute stays. Internal records reviewed by ProPublica showed Cigna doctors rejecting more than 300,000 payment claims through PxDx in two months in 2022, averaging 1.2 seconds apiece. Medicare's WISeR model now uses enhanced technology, including AI, to review selected services in six states; CMS says a licensed clinician must make a negative recommendation.8 A human signature does not decide whose standard the system serves.

Fragmentation is already visible too. Hospitals buy answer engines, record add-ons and specialised models one at a time. Each may help locally; together they amount to a procurement programme without a constitution. Payers, providers and patients then automate different sides of the same dispute. In 2024 insurers on HealthCare.gov denied 19% of in-network claims. Consumers appealed fewer than 1% of those denials, and insurers upheld 66% of the decisions that were appealed.9

Faster paperwork, same dispute. The centre figure is measured; the three-sided flow is a scenario. Automation can speed denials and appeals without settling who pays. KFF analysis of 2024 HealthCare.gov claims.

Payment denials differ from clinical errors, but they show how a consequential decision can be made at scale while the information and stamina needed to challenge it remain scarce. AI can let payers review faster and help providers and patients draft appeals faster, without settling who should pay. Each note, denial and appeal also enlarges somebody's store of data. Fragmentation therefore feeds consolidation: the organisation with the broadest view can learn faster than those supplying it.

Control changes clinicians as well as decisions. Aviation's safety engineers learned that a steep hierarchy may silence the person who spots a mistake. Automation researchers have long warned that turning a skilled operator into a monitor creates new failures. A doctor repeatedly asked to approve a machine's answer may become less able, or less willing, to reject it.

The machine at the shoulder. One observational study found worse unaided detection after routine AI use; one randomised trial found worse reasoning with flawed advice. Lancet Gastroenterol. Hepatol. 2025; NEJM AI 2026.

One observational study found that experienced endoscopists' unaided adenoma-detection rate fell from 28.4% before routine AI use to 22.4% afterwards; the design cannot establish that AI caused the fall. In the Pakistani trial, 20 hours of AI training did not prevent doctors from following plausible bad advice.10 “Human oversight” is a weak safeguard if the human has learned not to disagree. It describes a box in a diagram.

An institution with a memory

If oversight cannot be presumed, it has to be tested, and buying more tools will not do it. What settles the question is what happened to the patient and how the tool changed the work. Most studies never look. A review in March 2026 found that only 24% of 50 studies of predictive clinical decision support involved prospective use, and 64% reported technical measures without any data about workflow.11 A system may improve a score while learning little about care, and a good benchmark can conceal a bad clinical process.

Those harder studies need access to the records, and the record will be thinner where clinicians expect candour to be punished. The remedy therefore turns on ownership, which is still unsettled. One promising arrangement would combine inspectable software, shared control of the data and protected reporting. Open code permits scrutiny without opening a patient's file. Shared control leaves hospitals and public systems a say over how their data are compared and how tested improvements are reused. Reporting protections reward candour without excusing reckless or criminal conduct.

That last boundary is difficult. The prosecution of nurse RaDonda Vaught after a fatal medication error showed how quickly an argument about accountability can become an argument about whether anyone will report the next near-miss.12 A system that learns needs rules for blame, not a promise to abolish it.

Models and application code may become cheaper and easier to copy; a trusted institution will not. Its scarce assets are well-kept records, permission to compare them, the ability to test a change and a reputation for treating candour fairly. Those are products of governance, not properties of a model. Doctors need not own every model or every file. Aviation did not preserve the pilot by pretending automation could not fly. It gave pilots a new source of authority: stewardship of a system that learned from near-misses. Doctors have the same opening. If they do not take it, insurers, vendors and states will set the standard and collect the dividend. The machine will not heal itself. It has no stake in who does.