AI's deskilling question has a GI metric: ADR

Share

This article is part of GI & Hepatology News coverage of Cleveland Clinic's AI Summit for Healthcare Professionals, held Aug. 28 in Cleveland in collaboration with the College of Healthcare Information Management Executives. The summit is not specialty-specific, but gastroenterology sits at the leading edge of clinical AI adoption through computer-aided detection in colonoscopy, and questions raised there about implementation, expertise, and skill retention arrive in GI practices ahead of most other fields. GIHN is covering sessions with direct bearing on GI and hepatology practice.

 As artificial intelligence takes on more of the work of recognizing patterns in clinical data, the open question is what happens to the clinicians who used to do that work alone. Pete Clardy, MD, PhD, put it near the center of his keynote at Cleveland Clinic’s AI Summit for Healthcare Professionals, and said he could not answer it.

“I don’t have an answer,” he said. “But I do think that this is going to be one of the most important questions that we address going forward.”

For gastroenterology, the question is less abstract than it is for most specialties. Endoscopy was among the earliest fields to put AI in front of clinicians in real time, and it is one of the few with published data on what happens to performance when the software is switched off.

Dr. Clardy is a pulmonary and critical care physician and Director of Clinical Enterprise at Google for Health in New York, where he leads a team of clinicians working on scaling health care and life sciences innovation through emerging technologies and research.

Deskilling, misskilling, and never skilling

Dr. Clardy raised three related risks for clinicians, trainees, and medical educators as AI becomes more capable of storing information, organizing it, and judging its contextual relevance.

Deskilling describes the erosion of abilities a clinician once had. Misskilling describes the development of the wrong abilities. Never skilling, the category he flagged for learners who begin training in an AI-enabled environment, describes failing to develop the capacity to practice independently in the first place.

“How we manage [the widespread integration of AI] depends on where we are on our developmental curve,” he said.

What expertise should mean under those conditions, he suggested, remains unsettled.

What the colonoscopy data show

Endoscopy offers one of the few empirical tests of the first of those risks. In a multicenter observational study published in The Lancet Gastroenterology & Hepatology in 2025, Krzysztof Budzyń, MD, and colleagues examined adenoma detection rates at four Polish endoscopy centers participating in the ACCEPT trial, before and after those centers introduced computer-aided detection (CADe) for polyps at the end of 2021.

The analysis covered 1,443 patients who underwent standard, non-AI-assisted diagnostic colonoscopy during two three-month windows: 795 before the software was introduced and 648 afterward. ADR on those unassisted procedures fell from 28.4% to 22.4%, an absolute difference of −6.0% (95% CI, −10.5 to −1.6; P = .0089). In multivariable analysis, exposure to AI was independently associated with lower ADR (odds ratio, 0.69; 95% CI, 0.53-0.89).

The authors concluded that continuous exposure to AI might reduce the quality of standard colonoscopy, suggesting a negative effect on endoscopist behavior. The work was funded by the European Commission and the Japan Society for the Promotion of Science.

The design carries real limits. The study was retrospective and observational, the comparison windows were short, and the centers were concentrated in one country, so the findings cannot establish that AI exposure caused the decline. Patients with inflammatory bowel disease, prior colorectal resection, pregnancy, or intensive anticoagulant use were excluded. Subsequent correspondence in the journal argued that the training implications for novice endoscopists deserve particular attention, given that academic medical centers have been early adopters.

The finding also sits alongside a substantial literature pointing the other way while the tool is running. Comparisons of trainee, intermediate, and expert endoscopists have found that CADe can bring trainee ADR close to expert levels during AI-assisted procedures. The tension between those two bodies of evidence is the practical version of Dr. Clardy’s question: whether a skill supported by software is the same skill.

A ‘tool shaping moment’

Dr. Clardy described the present period as a window in which AI tools remain highly adaptable but will “harden over time,” which is why he argued that choices about their purpose matter now rather than later. He invoked Father John Culkin’s line that we shape our tools and thereafter our tools shape us.

He placed current systems earlier on that curve than the discourse often suggests. “AI is in its infancy,” he said. “It’s as bad as it will ever be right now, and the rate of change is remarkable.”

That evolution is arriving as clinicians face growing volumes of data. “We find ourselves collectively in this situation of too much data, not enough information,” he said.

From narrow tasks to co-clinicians

Dr. Clardy traced computerized decision support from early expert systems and supervised machine learning through large language models, agentic systems, and emerging “world models” designed to understand and simulate environments. Earlier systems relied on labeled data to learn narrow tasks, such as flagging abnormalities on imaging. Transformer-based systems can learn without supervision and produce more complex outputs, with recent work extending to frameworks in which multiple agents handle distinct tasks, as in Google’s Co-Scientist research.

“What we’re seeing … [is] helping to find signal from the noise, organize and summarize complex information, [but] not take over clinical decision-making,” he said. “But we are starting to see capabilities that go higher on that developmental pyramid.”

He described research aimed at expanding what an AI “co-clinician” can do, following a “hill climbing” approach modeled on how medical students learn to care for patients. In a single-center study he cited, a text-based system took a history of present illness, drew on the patient’s record, and produced documentation and a skeletonized assessment and plan. Patient trust increased after the interaction, and AI-generated differentials and management plans were similar in quality to those produced by humans. That work is moving to a multicenter study.

Other research extended the interaction to video, using a “talker” agent to maintain conversational flow and a “planner” agent to monitor for discrepancies and points needing clarification. Dr. Clardy said the system performed near human levels on triage and history taking and outperformed comparison models on clinical reasoning. Earlier text-based studies had identified an “empathy gap” favoring AI, while humans outperformed video AI on certain kinds of counseling and communication.

“This is really interesting research and also, yay humans,” he said.

The triadic relationship

Patients are adopting these tools quickly on their own, Dr. Clardy said, producing what he called an evolving “triadic relationship” among clinician, patient, and AI. He said the tools may help level the historic information asymmetry between patients and clinicians, allowing patients to review their records in detail and bring more informed questions to visits.

That trend has a specific texture in GI, where patients routinely receive pathology and endoscopy reports written in language the reports were never meant to explain, from dysplasia grading in Barrett’s esophagus to histologic activity scoring in inflammatory bowel disease.

Dr. Clardy tied the pace of patient adoption to a remark he heard at a National Academy of Medicine discussion, in which a patient advocate observed that innovation in health care moves at the speed of desperation.

“I think the dichotomy [between] trust and desperation comes when you think about how everyone needs an advocate and a way finder when it comes to managing their own health,” he said in an interview. “We in the health-care profession are learning a lot from the ways in which people are using these tools.”

Where implementation fails

Technological sophistication is not what determines whether an organization succeeds with AI, Dr. Clardy said. Success depends on change management, alignment, identifying the right stakeholders, and being “crystal clear” about the problem being solved, including whether the aim is to automate a process, augment human work, or pursue innovation. He encouraged starting with low-risk use cases.

“This is where we fail more often than on the basis of technology,” he said. “We are insufficiently crisp on the problem to be solved.”