How to Choose an AI Vendor in Behavioral Health

Behavioral health leaders sitting on the NatCon 2026 panel "Closing the Technology Gap: How AI Is Redefining Behavioral Health" had a warning for the room: choosing an AI vendor in 2026 is nothing like choosing an EHR in 2015. The technology moves faster, the risks compound faster, and the wrong choice can quietly hand your patient data to a company that turns around and uses it to train the next commercial model.
The good news, according to the panel, is that the questions to ask are actually knowable.
Three things every behavioral health leader should know
- Do not buy from any vendor that cannot tell you exactly what it does with your data. If they take PHI to train their models, you are trading value in a way you may not recognize.
- Do not fall for the "we built our own LLM" pitch. The frontier models have hundreds of millions of dollars behind them. Homegrown models fall behind within months.
- You need an AI governance council inside your organization, even a small one. Regulators cannot audit what you cannot inventory.
The State of AI in Behavioral Health: A One-Year Snapshot
Josh Sheller, CEO of Qualifacts, opened his section of the panel with a data point that stuck with the room. When his company launched AI-powered clinical documentation at NatCon 2024, roughly three-quarters of behavioral health agencies said they were not interested. One year later, at NatCon 2025, three-quarters of agencies were asking how fast they could implement it.
"Our clinicians want it. They're burning out, and they're trying to do more with less, and there's too much administrative overhead on them, and we needed to leverage this technology to make their job more efficient."
— Josh Sheller, CEO, Qualifacts
Sheller's own AI-generated estimate for the field, cross-checked against other sources: roughly 30% of U.S. behavioral health agencies had fully adopted at least one AI use case by the start of 2026. Another 25% planned to adopt AI in the same year, which would push the total above 50%.
That is only the integrated adoption number. If you count "shadow AI," the individual clinicians and administrators using ChatGPT or Gemini on their own devices, the real adoption rate is well above 90%. Which is exactly why governance matters.
The First Filter: What Does the Vendor Do With Your Data?
Josh Sheller was blunt about the risk. A lot of AI companies, he said, do not understand the responsibility that comes with PHI, and the integration points they use to reach it are often wide open.
"There's a lot of insecure connection points where companies are coming in through your user credentials, and the EHR has no control over what data is being extracted, nor no control over what data is being written back. For me, personally, as an EHR leader, [that] is a very scary piece of this."
— Josh Sheller, CEO, Qualifacts
Ryan Bach, CIO of Centerstone (the largest nonprofit behavioral health group in the country), was the operational voice on the same point: know where your data lives. Bach's approach at Centerstone is to go EHR-native first, keeping data inside the legal medical record wherever possible, and to avoid tools that would "farm our data off" to a third party. When his team has to bring in an outside vendor, they treat the data question as a hard filter, not a soft preference.
Sheller went further. There are, he said, a growing number of AI companies whose business model is not primarily software revenue. It is data acquisition.
"There's companies out there right now that are buying data to train models for tens of millions of dollars, but most of these companies are just going to unsuspecting agencies that have data, and having you pay them for their software, and they're getting your data for free to build new models."
— Josh Sheller
The industry, in Sheller's words, needs to be educated on this trade. Even when the software feels useful, the exchange of value may be dramatically asymmetric.
The February 2026 HIPAA & AI compliance analysis noted that the proposed HIPAA Security Rule update, expected to be finalized in May 2026, eliminates the "addressable" safeguard distinction and mandates annual risk assessments that must explicitly include AI systems. Business Associate Agreements now need language covering model output ownership, whether derivative PHI is generated, and which de-identification standard applies. If your vendor cannot answer those questions in writing, that is your answer.
Why Homegrown LLMs Are Losing
One of the most frequently marketed differentiators from smaller AI startups is the phrase "we built our own LLM." The panel was unanimous on how to interpret this claim.
The McKinsey QuantumBlack co-founder on the panel, now advising Warburg Pincus on healthcare investments, laid out the math. Frontier model developers such as OpenAI, Anthropic, and Google have collectively raised hundreds of billions of dollars to build and train their models. The cost per token for those frontier models is on track to fall roughly 900-fold, and it keeps dropping. No behavioral health-specific vendor, no matter how well-funded, can match that pace.
Sheller confirmed from Qualifacts' own testing. Homegrown LLMs have been compared side by side against off-the-shelf frontier models, and the frontier models are advancing so much faster that any bespoke advantage evaporates within months.
The panel's guidance: instead of buying a vendor with a proprietary model, buy a vendor with model flexibility. Ask whether their architecture can swap out an underlying LLM when a better one is released.
"What is their flexibility to leverage different LLMs and to continue to advance with the technology? Otherwise, you're going to be signing bigger contracts, [and] before your contract's up, your technology is going to be obsolete."
— Josh Sheller
Compliance Is a Mess. Governance Is How You Survive It.
Mike DeKock, the founder of MJD Advisors, is the panel's compliance and audit voice. His framing of the current regulatory landscape was blunt.
"AI regulations and ethics are, quite frankly, a mess. New proposals popping up at a state and even a city level. Trying to keep track of them all is a real challenge, but they all overlap pretty considerably, and they aren't actually all that restrictive."
— Mike DeKock, Founder & CEO, MJD Advisors
The problem is not that regulators are being overly aggressive. It is that they cannot enforce standards on activity you have not documented. Which is why Dec's core recommendation is procedural, not technical.
Every behavioral health organization, in Dec's view, should establish an AI governance council or working group. The council does not need to be large. What it needs to do is maintain a working inventory of what AI tools are being used, by whom, for what use cases, with what data, and with what human oversight checkpoints.
"The organizations that do this the best are going to be able to tell their story to customers of how they're using AI, how they're auditing their own use of AI, and how they're deploying it safely. I think that becomes a real differentiator in the marketplace."
— Mike DeKock
The Human-in-the-Loop Requirement
Every panelist returned to a shared premise: the models are getting better, but they are not yet good enough to make clinical decisions without human review.
Dec noted that in his non-healthcare consulting practice, some clients are beginning to ask about autonomous updates in production. In healthcare, that conversation is not close to being credible. What matters instead is where the human checkpoints live in the workflow, and whether they are documented well enough for an auditor to verify.
The panel also raised the growing risk of models trained to please. Chai flagged sycophancy — the tendency of models to tell users what they want to hear — as a specific concern, and the moderator pointed to growing research on AI-induced psychosis. Frontier models are optimized to give users the response they want to hear, which can be actively harmful when the user is in crisis.
"People were paying $200 a month to a company because, instead of getting a therapist, they had an AI life coach. The program was very reaffirming: 'No, you're good, you're great, you're doing a good job.' A lot of these people were in a really dark place, and they started having higher suicide rates because these people were not getting good therapy. They were getting affirmations of not healthy thoughts through the AI."
— Josh Sheller
A June 2026 industry compliance analysis made a related point: HIPAA compliance is necessary but not sufficient. HIPAA does not cover algorithmic bias, sycophancy, prompt injection, or the risks unique to generative systems. Governance frameworks have to go further.
A Practical Vendor Due-Diligence Checklist
Consolidating what the panel emphasized:
- Validate the company. Ask for referenceable customers. Confirm the leadership team, funding, and history. Basic vendor hygiene applies.
- Get every data question in writing. Where does the data live? Who accesses it? Is it used to train the vendor's models or anyone else's? What is the data-flow diagram?
- Confirm BAA language covers AI specifically. Model output ownership. Derivative PHI. De-identification standard (Safe Harbor vs. Expert Determination). Zero-retention configuration if using an API.
- Check for ISO 42001 certification (the international standard for AI management systems) or equivalent independent attestation.
- Ask which LLM(s) the vendor uses today, and how they will migrate when better models come out.
- Confirm human-in-the-loop checkpoints for every clinical use case, documented in a way that an auditor can verify.
- Ask about SOC 2 and HITRUST reports. Then ask who performed the audits. Compliance-as-a-service firms using AI to auto-generate audit reports are proliferating, and not all are equally credible.
- Verify who owns the data if the relationship ends. Portability language matters more than most contracts make it feel.
Frequently Asked Questions From the Audience
Q: The AI ethics questions worry us. How irresponsible does AI have to get before we should stop using it?
Sheller's answer: the marginal error rate for a well-designed AI system is often lower than the human error rate we already accept. The problem is that humans have a cultural tolerance for human error we have not yet extended to AI. The Tesla autopilot case that drove into a wall and killed someone got national headlines even though thousands of human-caused accidents happen every day. That scrutiny is uncomfortable, but Sheller argues it is also what drives AI to keep improving. The organizations that use AI recklessly, without oversight, are the real risk, not the technology.
Q: How much do we need to worry about garbage-in, garbage-out?
Chai gave the room a memorable insider anecdote. When OpenAI first trained GPT-3.5, model performance was 50% worse than expected because HTML tags had not been stripped from the training data. Simple data cleaning brought performance up dramatically. The takeaway: as you design any AI project, invest in data preparation before you invest in the model.
Q: What is Ryan most excited about in the next 12 to 18 months?
Bach's answer was personal. He lost a close friend and former boss to suicide in 2015. He is hopeful that the same recommendation engines that show us doomscroll content today can be redirected to serve as "conduits" to appropriate mental health resources at the moments people are most vulnerable, particularly at 2 a.m. when they cannot sleep.
The Bottom Line
The panel closed with a shared caution and a shared invitation. AI in behavioral health is not optional at this point. The wave is here, and the field is going to be shaped by whoever moves thoughtfully. What separates the organizations that will thrive from the ones that will get burned is not adoption speed. It is data governance, vendor due diligence, and the discipline of asking the right questions before signing the contract.