When Should an AI Agent Escalate to a Person?
Good AI knows what it doesn't know. Here are the escalation triggers worth wiring in.
The best AI agents aren't the ones that resolve the most. They're the ones that know exactly when to stop trying and get a person. An agent that pushes forward on a conversation it's losing does more damage than one that hands off early — and the customer remembers the difference.
So the real design question isn't "how do we make the AI resolve more?" It's "what should make the AI escalate?"
AI escalation starts with the agent knowing its limits
Good AI escalation rests on one idea: the agent should know what it doesn't know. A model that's confidently wrong is more dangerous than one that says "I'm not sure, let me get someone." Wiring escalation triggers in is how you turn that humility into behavior.
There are four signals worth treating as hard escalation triggers. Each one, when it fires, should route the conversation to a person — with full context attached so the rep continues instead of restarting.
The four triggers worth wiring in
1. Low confidence. When the agent isn't sure of its answer or the right action, it should hand off rather than guess. This is the foundation. Set a confidence threshold and respect it — under the line, the agent stops and gets a person. Fail closed, not open.
2. Sensitive intent. Some topics should go to a human regardless of how confident the AI is: billing disputes, cancellations, complaints about a person on your team, anything legal, anything involving money above your set limits. The agent's job here is to recognize the sensitive intent and route it, not to handle it.
3. Repeated failure. If the agent has tried twice and the customer is still stuck, the third try usually makes it worse. Count the attempts. After a small number of failed resolution attempts on the same issue, escalate. A loop is a tell that the agent is out of its depth.
4. Emotional cues. Frustration, anger, distress — these are signals that the customer needs a person, not a faster bot. The agent should read the tone and, when a customer is clearly upset, prioritize a warm handoff over one more automated attempt.
| Trigger | What it looks like | Right response |
|---|---|---|
| Low confidence | Agent unsure of answer/action | Hand off with what it knows |
| Sensitive intent | Refund dispute, cancellation, legal | Route to a person immediately |
| Repeated failure | Same issue, 2+ failed attempts | Stop, escalate with attempts logged |
| Emotional cues | Frustration, anger, distress | Warm handoff, fast |
An agent that hands off early with good context beats one that keeps trying and makes things worse. Knowing when to quit is a feature.
Don't just escalate — escalate well
A trigger that dumps the customer onto a cold queue isn't much better than no trigger at all. When any of these fire, the handoff has to carry context: the full history, the agent's read on intent, what it already tried, and the live account state. The customer should never have to repeat themselves at the moment they're already frustrated.
A few patterns that make escalation land:
- **Escalate the conversation, not just the customer.** The person should pick up mid-story, not from zero.
- Log the attempts. The rep needs to know what the agent already tried so they don't repeat the dead ends.
- Hand off fast on emotion. For an upset customer, speed of getting a human matters more than squeezing out one more automated try.
- Keep the receipt. The escalation itself is an action — record why the agent handed off, so you can review whether the trigger fired correctly.
Tune the triggers with real conversations
Escalation thresholds aren't set-and-forget. Too aggressive and you've built an expensive router that hands everything to humans. Too loose and the agent grinds through conversations it should have released. The way to find the line is to look at real conversations:
- Where the agent escalated, was it right to? Or could it have safely resolved?
- Where it kept going, should it have escalated sooner? Look for the loops and the frustrated customers it didn't release.
- Are sensitive intents being caught reliably, or is the agent occasionally trying to handle a dispute it should have routed?
Scoring every conversation — escalated and resolved alike — turns this from guesswork into tuning. You can see the false holds and the missed handoffs, and move the thresholds with evidence instead of instinct.
Escalation isn't a failure of the AI. It's the AI doing its job — knowing its limits and handing the hard, human conversations to the people who are good at them. Wire in the four triggers, make the handoff carry context, and tune the thresholds on real data.
Where BearScope fits
BearScope's AI agents escalate on low confidence, sensitive intent, repeated failure, and emotional cues — and every handoff carries the full thread, the agent's read on intent, and what it tried, with a receipt you can audit. Because every conversation is scored, you can see whether each escalation was the right call and tune from there. See the product overview, read how we handle audit and security, or book a walkthrough.
See it on your own conversations.
Bring your busiest day. We'll score every conversation in it.
Book a walkthrough →