Most health care leaders can tell you why they are adopting AI. Far fewer can show you how their governance actually runs at scale on a random Tuesday afternoon.
AI governance today is policy-heavy and operationally thin. Nearly every organization can hand you a polished governance document; very few can prove a specific output was reviewed, by whom and with what authority to stop it.
Accreditation closes that gap because it does not ask what you believe. It asks for evidence.
RediMinds recently became one of the first organizations to earn URAC’s Health Care AI Accreditation, the only comprehensive stand-alone accreditation program dedicated to AI in health care. It examines governance, defined uses of AI, risk management, transparency, monitoring and accountability across the lifecycle. Rather than evaluating principles; it evaluates whether they are visible in your workflows and production systems. We process hundreds of thousands of determinations a year, so this was never academic for us.
What I would tell any executive team considering the same process: trustworthiness is not a property of the model. It is a property of the foundation you build around it. No AI system, however accurate, is ready to scale where a determination affects a patient’s care or a clinician’s payment until a named human with the authority to stop it stands behind the output.
The person affected wears two hats. In utilization review, a member’s coverage determination decides whether a patient gets the care their clinician ordered. Under the No Surprises Act, the patient is deliberately held out of the fight, and the question is whether a clinician is paid fairly for care already delivered. Different stakes, same requirement.
Accuracy is not safety
The most expensive misconception in health care AI is that a high-performing model is automatically a safe one. Technical performance and safe, compliant use in the real world are different properties, and one does not deliver the other.
The risk lies in how the system reasons and what it is being asked to decide. A rules engine sorting invoices and a generative model summarizing a clinical record carry different failure modes, and the same model moved into a higher-stakes workflow becomes a different risk overnight. Treating all AI as one risk category means over-governing what is harmless and under-governing what is not.
Start with what is actually happening
A real AI inventory is not a list of vendors. For every use case, it captures the purpose, the data it touches, the decisions it influences, its risk tier, its named owner, the oversight required and how performance is monitored. That exercise exposes important gaps, like a vendor whose documentation is strong but whose tool was never validated on your population or a policy that assigns oversight broadly while giving no one the authority to pause a use case. Shadow AI surfaces here too, and is a good signal for what is missing. Employees reach for generative AI when approved options fall short, and a ban reduces reported use, not actual use.
Accountability means someone can stop it
“Human in the loop” has become a slogan. In practice, it has to mean something specific: a named accountable party with override authority at the point where the output has consequence for a patient or a member.
At our volume, a human cannot review every AI output; that would defeat the purpose of using AI. In our design, we need to decide where human judgment sits. We use tiered oversight, concentrating review where being wrong costs the most, and we document the override at the case level so any decision can be reconstructed.
Decide in advance who can approve, limit, suspend or retire a system and who accepts residual risk. If an incident happens, the first hours should not be spent establishing who is responsible. Training has to match: role-based and scenario-driven, covering when judgment overrides the tool, not annual and generic.
A pressure test, not a finish line
The most valuable output of accreditation was not the seal. It was a shared language across clinical, legal, compliance and operational teams, and a clear view of which processes were repeatable versus dependent on one person’s memory. An external reviewer working from a defined standard finds things an internal audit will not. But models drift and vendors update products between reviews, so monitoring and revalidation have to be normal operations, not a one-time project.
Six questions worth asking now
Do not wait for a regulator or an adverse event to ask for proof:
- Do we know everywhere AI is being used?
- Is one named person accountable for each use case?
- Can we show how each use case’s risk was evaluated?
- Are employees trained for the decisions they actually face?
- Are our disclosures to patients and members understandable and appropriate?
- Can we detect a problem and stop the tool quickly?
If those answers are not known by the leaders and the people using these tools daily, the model is not durable yet.
We pursued accreditation precisely because the regulatory landscape is uncertain. Federal and state frameworks are still forming, but they converge on the same signals: transparency, auditability, human oversight and documented attention to bias. Building for those principles and having a credible third party verify that you did so is more defensible now rather than retrofitting later.
AI will keep getting faster, and it should. But when a determination decides whether a patient receives care or whether a clinician is paid for care already delivered, a qualified human should stand behind it, equipped with AI and never replaced by it.
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.