Healthcare’s rush into generative AI has created a cost center many organizations only notice after the bill arrives: the token, the basic unit AI platforms use to bill for every prompt sent and every response generated.
As adoption scales from pilots to enterprisewide rollouts, health system leaders say spending can climb quickly, largely because the employees using these tools have no idea a meter is running.
“If you don’t manage this well, you will have runaway costs, because people don’t understand the cost,” said Luis Taveras, PhD, executive vice president and chief digital and information officer at Philadelphia-based Jefferson Health.
That gap in understanding is why Dr. Taveras and CIOs at several other health systems said the biggest lever they’re pulling isn’t a cheaper contract with an AI vendor — it’s discipline over who gets access to which tool, and for what.
At Jefferson, that discipline starts with a ladder. Employees with routine questions are told to stick with Google. Those who need more are pointed toward Jefferson’s basic Microsoft Copilot license, bundled into its existing Microsoft agreement at no extra cost, and then to Copilot Plus, a paid tier billed as a fixed monthly fee. Only employees who need more firepower than that get access to Anthropic’s Claude, or other frontier models billed by the token.
Dr. Taveras pointed to a recent example of what’s at stake in getting that ladder right. A colleague needed to analyze a large batch of files and asked for access to Cowork, Anthropic’s agentic desktop app. Jefferson granted it that morning, and by the end of the day, the employee had finished work that otherwise would have taken about two weeks — at a token cost of $75.
“If I think about this, a highly compensated person — two weeks’ worth of work in four hours — that’s not a bad return,” Dr. Taveras said. “But if he does that every day, the number is going to get pretty high.”
Jefferson has started allocating tokens by user type rather than giving everyone the same amount, tracking heavier users, such as its investment analytics team, differently from occasional users, and sending people reports on their own usage.
Certain tools are unlocked only after employees go through an education process on what the tools cost. That includes teaching employees to recognize how AI platforms are financially incentivized to keep a conversation going.
“The other thing we have to be careful with is what I call controlling the verbosity of these models. They’re very verbose — you ask a simple question, it comes down and gives you an answer, and then it says, ‘Can I do this? Can I do that for you?’ It keeps suggesting things to do,” Dr. Taveras said. “If you’re using one where the meter is running, it’s self-serving for the model to keep asking you for more, because you pay a lot more for what comes down than what you pay for what you send up. It’s going to keep suggesting more because it’s going to generate more money for the people running it.”
Jefferson is also navigating the enterprise side of token costs directly. “We’re negotiating with Anthropic now to make an investment in tokens,” Dr. Taveras said. “Once we make that investment, I’m going to really manage that token usage, and we’re going to look for the high-return areas, and that’s where we’re going to allocate those tokens.”
Charlottesville, Va.-based UVA Health is also being deliberate about which employees get access to which AI capabilities, rather than opening every tool to everyone at once.
“We fully recognize that there is tremendous interest in AI tools across UVA Health, and I support expanding access where it can meaningfully improve productivity, reduce administrative burden and make our employees’ lives easier,” said CIO Sonney Sapra. “However, what is often less understood is that every interaction with an AI platform consumes tokens, and as usage scales across a large organization, those costs can grow very quickly.”
“Our focus at UVA Health is to balance innovation with stewardship,” he said. “We are evaluating the appropriate levels of access, matching tools to specific business and clinical needs, measuring value and outcomes, and understanding utilization patterns before expanding broadly.”
At Atlanta-based Emory Healthcare, Chief AI Officer Nabile Safdar, MD, said his team applies a similar filter before a tool is turned loose enterprisewide.
“One of the primary levers we use to manage those costs is selecting the lowest-intensity model that is appropriate for the task at hand,” he said. Emory also requires AI vendors to open up their financial operations data before signing on, he said, “so we can better understand and forecast potential token consumption and associated costs.”
Eric Kirkendall, MD, chief medical information officer and CIO for academic health at Atrium Health Wake Forest Baptist in Winston-Salem, N.C., said the model-matching question has become a standard part of governance conversations there — not “what’s the best model,” but what’s the least expensive one that reliably gets the job done.
“Many organizations are still focused on reducing the cost of individual prompts. While that matters, we’ve found that the larger opportunity is ensuring AI is applied to the right problems, with the right model, and with clear measures of value before usage scales,” he said.
Several CIOs said that framing — value over price — is where the conversation is headed next, once basic tiering is in place.
“A high token bill may be entirely justified if the workflow reduces clinician burden, improves patient engagement, accelerates research or eliminates manual administrative work,” Dr. Kirkendall said. “We’re starting our journey to focus less on cost per token and more on cost per outcome delivered.”
At Palo Alto, Calif.-based Stanford Health Care, Chief Information and Digital Officer Michael Pfeffer, MD, said cost enters the conversation before a project is even greenlit.
“We start with a deep understanding of the problem we’re trying to solve to determine the best possible IT solution, which may or may not be AI,” he said. “If it does require AI, then we use our FURM assessment framework to understand the costs relative to the value of the solution before proceeding. We monitor all of our AI tools once live in production for system integrity, performance and impact.”
At Houston Methodist, Chief Innovation Officer Roberta Schwartz, PhD, said a standing internal committee vets proposed AI use cases for return on investment before the system decides whether to build, buy or partner on a given tool, then keeps watching after it goes live.
“Like any technology we implement, agentic AI is subject to ongoing governance, monitoring and optimization to ensure it continues to perform as intended and deliver value safely and effectively,” she said.
Adam Landman, MD, chief digital information officer at Providence, R.I.-based Brown University Health, takes a different tack on exposure: pricing structure itself. Most of the AI capabilities Brown has deployed at scale run on flat per-user pricing rather than token-based consumption, he said, which shields the system from utilization surprises. That won’t hold for every tool, he added, particularly the most powerful task-executing agents, where the underlying compute is genuinely expensive.
“For consumption-priced capabilities, we start with a small pilot group and usage limits rather than broad access. That caps our exposure while we learn, and it generates real utilization data so we can forecast annual cost with confidence before scaling,” Dr. Landman said in a written statement. “Underneath all of it, the enabling capability is FinOps. Real-time visibility into consumption and spend is essential — if costs move unexpectedly, we want to know within hours to days, so we can act while the number is still small.”
Dr. Landman also pointed to a longer-term shift some organizations are beginning to weigh: moving away from per-token vendor pricing altogether.
“Token-based pricing is challenging for any organization because cost scales with adoption, and adoption is the goal,” he said. “That’s one reason many organizations are exploring ‘AI sovereignty’ — running open-weight models on dedicated infrastructure, where cost is driven primarily by the hardware investment rather than each individual transaction.”
For now, most of the leaders interviewed agreed that the winners of this phase of AI adoption won’t be the systems that land the cheapest per-token rate.
“The organizations that will manage AI costs most effectively won’t necessarily be those with the cheapest tokens,” Dr. Kirkendall said. “They’ll be the ones with the strongest governance, the best reuse of established capabilities and the clearest understanding of where AI creates measurable value.”
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.