Sponsored

When every answer sounds right: Why clinical intelligence needs a validation layer

Advertisement

For most of medicine’s history, the evidence a clinician needed was scarce, scattered and difficult to retrieve at the point of care. Today, the problem has shifted. Clinicians are no longer practicing in an environment defined by too little information. They are practicing in one defined by abundance: clinical studies, guidelines, summaries, search results and now, AI-generated answers that appear instantly, sound authoritative and are only a search bar away.

That abundance places new demands on clinical judgment. “The era has shifted,” said Sheila Bond, MD, Director of Clinical Content Strategy at Wolters Kluwer Health. “Access is no longer the problem. It’s what you do with what you can access.” For Dr. Bond, this is the practical challenge of clinical AI. The issue is not whether generative AI can produce fluent answers. It can. The question is whether those answers are grounded, valid, clinically appropriate and accountable to the standards medicine has spent decades building.

The skill that matters now is discernment: knowing which answer to trust when a patient is in front of a clinician. That is why clinical intelligence, the layer that pairs expert, evidence-based knowledge with machine synthesis, is only as trustworthy as the validation layer beneath it.

The risks of unvalidated AI

Dr. Bond is deliberate about the language of trust. “Clinicians should not be asked to trust an AI tool in the same way they might trust a seasoned colleague,” she said. Instead, they should apply the same critical discipline that has long defined evidence-based medicine.

“You come to everything with a critical eye, that’s what our field is about,” Dr. Bond said. “We teach critical appraisal and evidence-based medicine.” Generative AI becomes a new object for that same discipline. Clinicians need to ask where the answer came from, what evidence supports it, whether it applies to the patient in front of them and what limitations may be hidden beneath a confident tone.

Rather than placing trust in the model itself, Dr. Bond argues that clinicians and health systems need confidence in the governance around it. That means understanding who is accountable for the system, how information flows through it, how outputs are measured and how the tool is validated before and after it reaches clinical use.  AI excels at synthesizing. It can organize information quickly, summarize complex material and produce language that feels coherent. But synthesis is not the same as independent clinical reasoning within a specific domain. In medicine, that distinction matters.

When an AI tool reaches the point of care without adequate validation, Dr. Bond sees two distinct risks. The first is immediate and practical. “I’ve been in medicine for over 20 years, and I’ve seen how words are used,” she said. “When guidance comes from a source that sounds authoritative, and decisions are rushed, clinicians may act on it exactly as presented.”

The second risk is quieter but potentially more consequential: the gradual erosion of the critical appraisal culture that evidence-based medicine depends on. Evidence-based medicine was built to prevent this kind of shortcut. Its purpose is to separate what is merely plausible from what is clinically sound through disciplined appraisal, synthesis, review and consensus. “Evidence-based medicine is not one person looking at one study and deciding what it means,” Dr. Bond said. “It requires diversity of perspective, depth of knowledge and methodologic rigor to keep us honest.”

The danger, she said, is ceding that discipline to a model that can generate an answer that is pleasing, plausible and fluent. But plausible is not the same as valid. Without a critical eye and without a validation layer, the difference can become difficult to discern.

Beyond the benchmark

Much of the public conversation about clinical AI trust has centered on benchmarks: medical exam scores, case challenges, leaderboard performance and other single-point measures. Dr. Bond welcomes the scrutiny but cautions against overinterpreting any one metric. Medicine already understands the limits of benchmarks.  Licensing exams, board scores and quality ratings all provide useful information. None of them captures the totality of clinical judgment. “Those single data points don’t encapsulate who will become a good physician, or the totality of what being a good clinician is,” Dr. Bond said. “The same goes for AI models.”

For Dr. Bond, evaluating something genuinely new in medicine requires a multidimensional approach. A benchmark may show that a model can perform well on a constrained task. It does not necessarily show that the system can behave reliably in clinical practice, remain grounded in trusted knowledge, handle ambiguity or fail safely.

At Wolters Kluwer, the validation layer behind UpToDate Expert AI is designed around four dimensions.  The first is clinical intent: does the system do what clinicians need and expect it to do? That requires both human expert review and predefined automated measures. Wolters Kluwer draws on more than 7,600 contributors and tens of thousands of criteria that define, in advance, what “good” looks like.

The second is knowledge integrity, a dimension Dr. Bond believes is often underappreciated. The goal is to ensure that answers are traceable to authoritative content and that the system is not making unsupported leaps from isolated studies or single clinical claims.

The third is deliberate stress testing for risk and edge cases through red teaming. As part of that process, Dr. Bond said, teams continuously challenge the system, attempt to provoke errors and identify failure modes before they reach users.

The fourth is a continuous learning loop, in which clinician feedback is audited and fed back into both the system and the underlying content. Validation is not a one-time event. It is an ongoing responsibility.

What leaders should ask

For clinical and health IT leaders evaluating AI tools, Dr. Bond emphasizes diligence and accountability:

  • Who is responsible for the system?
  • What principles guide its development?
  • What governance processes oversee it?

From there, leaders should examine how knowledge flows through the system.

  • What sources does the tool rely on?
  • How are those sources curated?
  • Can answers be traced back to trusted content?
  • How often does the model use the intended knowledge base?
  • What are its hallucination rates?
  • Has the organization rigorously studied how the tool is likely to fail?

“I don’t make diagnostic tests, but I use them and I know their performance characteristics,” Dr. Bond said. “It’s the same with AI. You should understand how well your AI works and what makes it trustworthy.”

Clinicians cannot be removed from that process. The people who know what’s worth measuring are the people closest to patient care. Many AI benchmarks have emerged from the technology sector, but Dr. Bond argues that the meaning of what is measured must come from the clinical perspective.

AI in healthcare should be judged the same way other technologies introduced into patient care are judged: not by how sophisticated its components are, but by whether it supports better decisions, safer care and stronger outcomes for patients.

Generative AI can be made to move fast. Getting it right for healthcare is both more difficult and more important. Like any technology used in patient care, it should be held to rigorous standards: grounded in trusted sources, transparent about how it produces its outputs and continuously validated by the clinicians and organizations accountable for its use.

“This is an epistemic shift,” Dr. Bond said. “It’s going to change how we think and reason. It’s going to change us. So we’ve got to do it with extreme care.”


At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.

Advertisement

Next Up in Artificial Intelligence

Advertisement