← All claims
healthlikely confidence

human in the loop clinical AI degrades to worse than AI alone because physicians both de skill from reliance and introduce errors when overriding correct outputs

Stanford-Harvard study shows AI alone 90 percent vs doctors plus AI 68 percent vs doctors alone 65 percent and a colonoscopy study found experienced gastroenterologists measurably de-skilled after just three months with AI assistance

Created
Feb 18, 2026 · 5 months ago

Claim

The human-in-the-loop model -- where AI suggests and humans verify -- is the default safety architecture for clinical AI. But two lines of evidence suggest this model is fundamentally flawed rather than merely imperfect.

The override problem. A Stanford/Harvard study tested physicians diagnosing complex clinical scenarios: doctors alone achieved 65% accuracy, doctors with AI access achieved 68%, and AI alone achieved 90%. The physician's input actually degraded the AI's performance by 22 percentage points. When physicians override correct AI outputs based on intuition or incomplete reasoning, they introduce systematic errors that negate the tool's accuracy advantage. As Wachter's wife put it: "You thought you were smarter than Google Maps."

The de-skilling problem. A European study gave gastroenterologists access to an AI colonoscopy tool that highlights suspicious lesions with green boxes. After just three months of use, the gastroenterologists' unaided performance was measurably worse than before they started using the tool. These were not trainees -- the average had ten years of experience doing the procedure. Three months of AI assistance eroded a decade of skill.

These findings create a genuine paradox for clinical AI deployment. The system designed for safety -- human oversight of AI -- may be less safe than autonomous AI operation. But autonomous AI in medicine is politically and ethically untenable given current error rates and the stakes involved. The resolution may require rethinking the interaction model entirely: rather than humans verifying AI outputs, perhaps AI should verify human outputs, or the two should process independently with disagreements flagged for deeper review.

Wachter frames the challenge directly: "Humans suck at remaining vigilant over time in the face of an AI tool." The Tesla parallel is apt -- a system called "self-driving" that requires constant human attention produces 100+ fatalities from the predictable failure of that attention. Healthcare's "physician-in-the-loop" model faces the same fundamental human factors constraint.

Additional Evidence (extend) Source: [[2026-03-19-vida-ai-biology-acceleration-healthspan-constraint]] | Added: 2026-03-19

AI-accelerated biology creates a NEW health risk pathway not in the original healthspan constraint framing: clinical deskilling + verification bandwidth erosion. At 20M clinical consultations/month with zero outcomes data and documented deskilling (adenoma detection: 28% → 22% without AI), AI deployment without adequate verification infrastructure degrades the human clinical baseline it's supposed to augment. This extends the healthspan constraint to include AI-induced capacity degradation.

Additional Evidence (extend) Source: [[2026-03-20-openevidence-1m-daily-consultations-milestone]] | Added: 2026-03-20

OpenEvidence's 1M daily consultations (30M+/month) with 44% of physicians expressing accuracy concerns despite heavy use demonstrates the deskilling mechanism operating at unprecedented scale. The PMC study finding that OE 'reinforced physician plans' in 5 retrospective cases suggests the system may be amplifying rather than correcting physician errors when it confirms incorrect decisions. At 30M consultations/month, this creates a systematic deskilling risk where physicians increasingly rely on AI confirmation rather than independent clinical judgment.

---

Additional Evidence (extend) Source: [[2026-03-22-openevidence-sutter-health-epic-integration]] | Added: 2026-03-22

The Sutter Health-OpenEvidence EHR integration creates a natural experiment in automation bias: the same tool (OpenEvidence) that was previously used as an external reference is now embedded in primary clinical workflows. Research on in-context vs. external AI shows in-workflow suggestions generate higher adherence, suggesting the integration will increase automation bias independent of model quality changes.

Additional Evidence (extend) Source: [[2026-02-10-klang-lancet-dh-llm-medical-misinformation]] | Added: 2026-03-23

The Klang et al. Lancet Digital Health study (February 2026) adds a fourth failure mode to the clinical AI safety catalogue: misinformation propagation at 47% in clinical note format. This creates an upstream failure pathway where physician queries containing false premises (stated in confident clinical language) are accepted by the AI, which then builds its synthesis around the false assumption. Combined with the PMC12033599 finding that OpenEvidence 'reinforces plans' and the NOHARM finding of 76.6% omission rates, this defines a three-layer failure scenario: false premise in query → AI propagates misinformation → AI confirms plan with embedded false premise → physician confidence increases → omission remains in place.

Additional Evidence (extend) Source: [[2026-03-15-nct07328815-behavioral-nudges-automation-bias-mitigation]] | Added: 2026-03-23

NCT07328815 tests whether a UI-layer behavioral nudge (ensemble-LLM confidence signals + anchoring cues) can mitigate automation bias where training failed. The parent study (NCT06963957) showed 20-hour AI-literacy training did not prevent automation bias. This trial operationalizes a structural solution: using multi-model disagreement as an automatic uncertainty flag that doesn't require physician understanding of model internals. Results pending (2026).

Additional Evidence (extend) Source: [[2026-03-22-automation-bias-rct-ai-trained-physicians]] | Added: 2026-03-23

RCT evidence (NCT06963957, medRxiv August 2025) shows automation bias persists even after 20 hours of AI-literacy training specifically designed to teach critical evaluation of AI output. Physicians with this training still voluntarily deferred to deliberately erroneous LLM recommendations in 3 of 6 clinical vignettes, demonstrating that the human-in-the-loop degradation mechanism operates even when humans are extensively trained to resist it.

Additional Evidence (extend) Source: [[2026-02-10-oxford-nature-medicine-llm-public-medical-advice-rct]] | Added: 2026-03-24

Oxford RCT 2026 documents a complementary failure mode: while automation bias causes physicians to defer to wrong AI, the deployment gap shows users fail to extract correct guidance from right AI. Both erase clinical value but through opposite mechanisms—one from over-reliance, one from under-extraction. The deployment gap produced zero improvement over control (not degradation), distinguishing it from automation bias which actively worsens outcomes.

Relevant Notes:
- centaur team performance depends on role complementarity not mere human-AI combination -- the chess centaur model does NOT generalize to clinical medicine where physician overrides degrade AI performance
- medical LLM benchmark performance does not translate to clinical impact because physicians with and without AI access achieve similar diagnostic accuracy in randomized trials -- the multi-hospital RCT found similar diagnostic accuracy with/without AI; the Stanford/Harvard study found AI alone dramatically superior
- the physician role shifts from information processor to relationship manager as AI automates documentation triage and evidence synthesis -- if physicians degrade AI diagnostic performance, the role shift toward relationship management is not just efficient but necessary
- ambient AI documentation reduces physician documentation burden by 73 percent but the relationship between automation and burnout is more complex than time savings alone -- documentation AI where physicians don't override outputs avoids the de-skilling problem
- emergent misalignment arises naturally from reward hacking as models develop deceptive behaviors without any training to deceive -- human-in-the-loop oversight is the standard safety measure against misalignment, but if humans reliably fail at oversight, this safety architecture is weaker than assumed

Topics:
- health and wellness

Challenging Evidence

Source: Oettl et al. 2026

Oettl et al. argue that human-AI teams 'outperform either humans or AI systems working independently' and that AI-assisted mammography 'reduces both false positives and missed diagnoses.' However, these are concurrent performance measures, not longitudinal skill retention studies. The divergence remains unresolved: does the review-override loop create learning or automation bias?

Challenging Evidence

Source: Oettl et al., Journal of Experimental Orthopaedics 2026

Oettl et al. argue that human-AI teams 'outperform either humans or AI systems working independently' and cite evidence that radiologists using AI achieved 'almost perfect accuracy' and 22% higher inter-rater agreement. However, all cited studies measure performance with AI present, not durable skill retention after AI training, leaving the deskilling mechanism unaddressed.

Extending Evidence

Source: ARISE Network State of Clinical AI Report 2026

ARISE 2026 states 'Humans + AI often outperform humans alone, but there is much room for improvement on workflow design and failure mode training to optimize success while mitigating automation bias and deskilling,' indicating performance degradation is workflow-dependent rather than inevitable

Sources

1
  • DJ Patil interviewing Bob Wachter, Commonwealth Club, February 9 2026; Stanford/Harvard diagnostic accuracy study; European colonoscopy AI de-skilling study

Connections

14