TLDR
Build the top of the stack, source the layer underneath. Prompts, frameworks, and interface are fast to build and worth owning. The behavioral foundation beneath them, how each person on your team operates, how two specific people work together, and the standard that keeps advice pointed toward the growth of the relationship, could be the most important aspect to consider.
An HR leader pushed back on us recently, and the question was a good one: “Why would I use your platform if I can give ChatGPT or Gemini the same data and it spits out the same information?” It is exactly the right thing to ask. Building your own AI coach inside Claude, Copilot, or ChatGPT has never been more possible, and for the fast, exploratory part of the work, you should do it.
So this is not a piece about why you cannot build your own. You can, and plenty of sharp People teams already are. It is an honest look at what a home-built AI coach gets right, the five gaps that stall most of them, and the one part worth investing in instead of rebuilding.
This is not a decision you can put off, either, because your people are not waiting for a policy. About one in five U.S. workers now use AI on the job, up from 16 percent a year earlier (Pew), and roughly 45 percent use it at least a few times a year (Gallup). They are not only using it for tasks: the top reported use of generative AI is now therapy and companionship (HBR), and survey after survey finds employees would rather ask a chatbot for advice than their own manager, because it feels less judgmental (HR Dive, Fortune).
Whether or not you build a coach, your people are already using LLM’s to do so. The only question is what how great is it coaching them?
Why People and L&D teams are building their own AI coaches
The pressure is real. Boards and CEOs want proof of AI progress, and “we are building our own AI coaching” is a defensible answer. The tools finally make it feasible: an HR leader can stand up a coaching skill in Claude in an afternoon, load in the company competency model and values, and have something that talks.
In engineering-heavy cultures it is the house style. As one People leader told us, “a lot of our tech folks are vibe coding, that is very much the way of the world right now.”
The pull is not only bottom-up. Diane Penn, Anthropic’s first technical product manager, has said she wants her teams using AI “to have better conversations with each other, to be better managers,” and built a coaching skill of her own to prep for hard conversations.
When a product leader at a frontier lab is hand-building the prompt layer, it is fair to ask why you would not. There is even a real-world example on the record. On the Modern People Leader panel, Sarah Royer, who leads People Ops at Nirvana Insurance, walked through building her own AI coach as a skill in Claude, layered with her company’s competencies and her CEO’s message so the guidance stays consistent. She is candid about where it helped and where it stalled.
What are the benefits of building your own AI coach?
A do-it-yourself coach in an LLM can be genuinely good at a few things.
Speed and experimentation.
You can prototype a coaching prompt in an afternoon and change it the next day. No procurement cycle, no vendor timeline.
Your language, your frameworks.
You can load in your competency model, your values, and your leadership expectations so the coach speaks the way your company speaks.
Cheap to start.
If your org already pays for Claude, Copilot, or ChatGPT, the first version costs you time, not budget.
This is the right place to start, and a prototype can teach you what you actually need. The trouble can show up later, when that prototype has to become a tool various teams and people rely on every week.
We wanted to know where that ceiling is, so we tested it. In a study we call AI as a Workplace Coach, we ran five of the most widely used enterprise AI models through five common workplace conflicts, three times each, and had three coaches blind-score every response on the components of emotional intelligence, plus how specific the advice was.
The short version: on its best day the raw model is a mediocre relationship coach, and the same five gaps show up every time. The strongest model scored 16.5 out of 25, rated solid but incomplete, and none reached relationally intelligent.
See how AI is coaching your people at their most frustrated
Cloverleaf Labs ran five leading models through five real workplace conflicts to show whether the AI coaches toward repairing a relationship or toward protecting themselves against it.
5 risks of building your own AI coach
We saw the same five risks across the LLM Relational Coaching Study and dozens of customer conversations. None is a reason not to build. Each is a reason to go in knowing what you are taking on.
Active listening: understand before you respond
Most conflict is less about the problem itself and more about how people feel about it, and nothing escalates frustration faster than feeling unheard. Skilled managers paraphrase to confirm understanding, ask clarifying questions to surface the real concern, and acknowledge emotion before moving to solutions.
01
1. It runs on a model, not a behavioral foundation
An LLM plus your HR documents is a chatbot with HR documents. It is not coaching, because it does not know the people involved. Our Relational Work Index, a study of nearly 30,000 real coaching conversations, found that about half of all coaching questions name a specific colleague rather than a skill or concept. People do not ask “how do I give feedback”; they ask “how do I give feedback to this person.”
WHAT THE BENCHMARK FOUND
Whether the advice fit the actual people involved or could have gone to anyone was the weakest thing the models did, scoring 45 percent. It was also the one dimension that did not improve as the situation changed: the models kept producing detailed action plans no matter the power dynamic, they just stopped doing the relational work underneath.
A general-purpose model with no validated behavioral assessment data and no relationship-level context gives advice that sounds right but stays generic. One HR leader described her workaround honestly: “this is the poor man’s version of what you guys do, I throw all my assessments into Claude and start asking questions.” It helps her as an individual, and it hits a ceiling fast.
02
It agrees with whoever is asking
A model optimized to be agreeable tells you what you want to hear, and that is the opposite of good coaching. A researcher we work with put it plainly: be careful using raw LLMs for coaching, because a model built to be agreeable “tells you what you want to hear.” The stakes are higher than that sounds. Research on self-awareness finds about 95 percent of people believe they are self-aware while only 10 to 15 percent actually are (HBR), so a coach that emphatically reinforces a shaky self-read does not just miss, it erodes the person’s ability to stay curious and collaborate.
WHAT THE BENCHMARK FOUND
When the person was frustrated, the models validated first and informed second, and at the far end one met a frustrated employee with “your frustration is not just valid; it’s a sign that you are deeply sane.”
Two of the pillars this depends on, challenging how the person sees themselves and building an understanding of the other person, both averaged below the neutral midpoint of 3 out of 5.
When the people building the frontier models say the default is to agree and the hard part is teaching it not to, that agreeable default is the ground your build starts from.
03
It overwhelmingly coaches people to protect themselves, not repair the relationship
When coaching that touches a real relationship gets hard, the raw model reaches for self-protection: document it, build your case, keep the receipts.
WHAT THE BENCHMARK FOUND
No model reached the repair-oriented end of the scale on average, and not one of the five coached toward investing in the relationship.
Of 638 distinct pieces of advice the models gave across the 75 conversations, three coached the employee toward genuinely repairing the strained relationship, and the rest were about winning, surviving, or managing the situation. Across all 75 responses, 52 percent steered the employee toward self-protection and only 12 percent toward investing in the relationship.
In the scenario about a controlling boss, every model in every run, fifteen out of fifteen, raised leaving as the answer rather than working it out.
The pattern tracked power: when the employee held authority over the other person, relational scores averaged 3.5 out of 5, but when a boss held power over them those scores fell to 2.1, roughly a 40 percent drop.
At its worst a model did not just fall short, it took sides: in the reorg scenario one described the team as having split into “factions,” cast a coworker as “the opposing tribe,” and coached the person to rally her “allies.” Coaching that changes behavior does the opposite. It moves people toward each other, not into defensive corners, which is exactly the relational work AI is least equipped to do on its own.
04
It is a tool people have to remember to open
The adoption killer is that a home-built coach is usually a destination. Someone has to remember it exists, open it, and re-set the context every time. Coaching that changes behavior does the opposite. It arrives in the moment, in Slack, Teams, and the calendar, tied to the meeting or the person in front of you, so it costs the employee no extra effort. That is a delivery system, and it is a different thing to build than a model.
WHAT A BUYER HAS TOLD US
“I have to re-prompt and re-upload things every time, it is not a long-term tool.”
05
You own the legal risk and the upkeep
Two burdens come with the build that the estimate almost never covers. First, the risk: coaching that touches performance, promotion, or pay is legally loaded, and the benchmark showed a model will deliver relationally destructive advice in the same confident, reasonable voice it uses to summarize a report. Across 25 model-and-scenario tests, genuinely strong coaching appeared once, while advice bad enough to be scored sycophantic or adversarial appeared six times. Nothing flags the difference, so no one catches it, and accountability for the answer lands on whoever built the coach. Second, the upkeep: models change, things break, and someone has to maintain the prompts, refresh the content, and document how it all works.
WHAT WE HAVE SEEN
One company built its own agent on Claude, Gemini, and ChatGPT, then could not train anyone on it, “I cannot train on something when I do not know how it works.” Another built a standalone coaching chatbot, found it did not work, and switched it off.
A thinking partner doesn’t just agree with you. It should add to you.
Diane Penn
First technical product manager at Anthropic, speaking on Lenny's Podcast.
AI coaches need behavioral data and relational context
The parts that are fun and fast to build, your prompts, your frameworks, your interface, are the parts worth building. The part that is slow, expensive, and easy to underestimate is the behavioral foundation underneath, and that is the part worth buying. It is the exact thing the benchmark showed a raw model cannot reproduce.
13+ market leading assessments, synthesized
Cloverleaf combines 13+ market leading behavioral assessments into one read on how a person works, through a patented synthesis.
Four layers of context, at once
Who a person is, who they are working with, what is changing in the organization right now, and what the coaching has learned over time. Most tools hold only one of these. Users with all four active are 90x more engaged than people working with a general-purpose AI assistant.
Relationship-level intelligence.
Coaching for how two specific people work together, where they are likely to align and where they might clash, mapped across more than a million behavioral signals. This is exactly what the benchmark showed the raw model cannot see, and about half of coaching questions are about a specific colleague, so it is the layer that answers them.
Eight years of signal.
65 million coaching moments and two approved patents behind the synthesis and relationship engine. That corpus cannot be rebuilt from a prompt, and it compounds.
There is also a strategic reason not to over-invest in a bespoke build. As one leader put it, the foundation models are getting more powerful so quickly that “if you don’t figure out some way to direct that, you’re just going to be eaten.” The durable move is not to out-build Claude or Copilot. It is to direct them with a behavioral layer they cannot reproduce, and a relational standard they do not meet on their own.
How to add behavioral data to your AI coach with MCP
Cloverleaf offers an integration over the Model Context Protocol (MCP), the open standard AI assistants use to connect to trusted data. It lets you pull Cloverleaf’s behavioral layer into Claude, Copilot, ChatGPT, or the internal AI your company built, so the coach you are building has context on the actual people involved.
Cloverleaf is deliberately careful with data. Only the behavioral context relevant to the question is passed, never raw scores or HR records. Cloverleaf does not train external AI models on your workforce data. And for the sensitive, personal work, the private Cloverleaf platform stays the place for it, so people can be candid. It adds coaching to the AI you already have; it does not replace your platform or your build.
IN PRACTICE
List what you are planning to build and what you are planning to source. Most teams we work with build their own prompts and skills, and source the behavioral data, the assessment IP, and the delivery mechanism from Cloverleaf. The build no longer starts from zero.
WHAT TO BUILD
- Your prompts and skills
- Your competency frameworks and company language
- the interface your people use
- and the internal story of your AI strategy
WHAT TO SOURCE
- The validated assessment IP and the synthesis across them
- The relationship-level data between specific people
- The daily delivery into the flow of work
- The relational standard and guardrails around coaching advice
You own the build, your prompts, frameworks, and interface, and source the one layer a model can’t generate: the behavioral foundation and relationship context. And if you would rather not build at all, that is a fair choice too. Here is an honest look at the AI coaching platforms worth comparing.
FAQ: building your own AI coach
If I give ChatGPT or Gemini the same data, will it not give the same answer?
Not for coaching. A general-purpose model can summarize your documents, but it does not hold validated behavioral data or the relationship context between two specific people, and about half of coaching questions are about a specific colleague. When we tested it, whether the advice fit the actual people involved was the weakest thing the models did, at 45 percent. Without that layer, the advice stays generic. The model is the easy part; the behavioral foundation is what makes the answer specific.
Can I build my own AI coach in Claude or Copilot?
Yes, and it can be a good place to start. You can load your frameworks and prompts and prototype quickly. Where builds stall is the behavioral data foundation, the tendency to validate a frustrated user rather than coach them, delivery in the flow of work, and ongoing maintenance. Many teams build the prompt layer and source the rest.
Should we build our own AI coach, and would it be cheaper?
Initially, it can look cheaper because the first prototype mostly costs time. The real cost shows up in the parts the estimate skips: assessment IP, integrations, a relational standard for the advice it produces, delivery, and maintenance. Costing those honestly usually reframes it from build-versus-buy to build-and-buy.
Is a self-built AI coach compliant and safe?
That responsibility is yours. Coaching that touches performance, promotion, or pay is legally loaded, and our testing found a raw model will give confidently wrong or relationally harmful advice in the same reasonable voice it uses for everything else, so nothing flags it. A home-built coach needs vetted content, guardrails, and a clear owner for the advice it produces. This is one of the main reasons People teams source the coaching layer rather than build it alone.
Can I just upload our frameworks to LLM's and get a coach customized for us?
You can get something that sounds customized. But a coach that only knows your documents still does not know your people, and it will still default to validating whoever is asking. Customization that changes behavior comes from combining your frameworks with validated behavioral data and relationship context, which is the layer worth sourcing.
Build your AI coach on a behavioral data foundation, not from scratch
Build versus buy was never really the question. Your people are already taking their hardest moments at work to an AI, privately, at their most frustrated, when what it says back matters most. The only real question is whether it sends them back toward each other or away.
Left to the raw model, it is away. AI as a Workplace Coach found the models coach self-protection over repair, side with whoever is typing, and, when a boss is involved, point toward the door, all in a calm, reasonable voice that never flags the harm. That is the default your build starts from.
So build what is yours to build: the prompts, the frameworks, the interface. Source the one thing a model cannot generate on its own, a real read on the people involved and a standard that keeps advice pointed toward the relationship. Do that, and the coach your managers open on a Tuesday actually knows them and moves them toward each other, instead of teaching your best people to lawyer up and leave.
See what behavioral data adds to your AI coach.
SOURCES
Cloverleaf, AI as a Workplace Coach (5 leading AI models, 5 workplace-conflict scenarios, 3 reps each, 75 responses, blind-scored by 3 human coaches on self-awareness, accountability, other-awareness, relational repair, and specificity). Cloverleaf, Relational Work Index, 2026 (29,374 coaching conversations). Modern People Leader panel (Kirsten Moorefield; Sarah Royer, Nirvana; Sarika Lamont, Vidyard). Lenny’s Podcast (Diane Penn, Anthropic). Pew Research, October 2025. Gallup, Q3 2025. Harvard Business Review, 2025 and 2018. HR Dive. Fortune.