Insights
Does your customer support AI agent really know when to escalate?
By the end of the year, 88% of companies expect to have AI agents handling some of their customer conversations, according to our survey of 2,527 enterprise leaders across 10 countries.
That doesn’t mean human representatives are out of the picture. Gartner reported that nearly 80% of customer service organizations plan to reshape the roles of some of their human agents as routine work gets automated while high-stakes conversations still need a person.
When AI fails to help, the challenge is ensuring the switch from AI to human happens when it should, and that part has to be designed. Where nobody sets the right boundaries, the agent answers everything, at the brand’s own risk.
When AI agents fail to help
AI agent failures in customer communications make the news regularly. Sometimes, the cause boils down to the agent having no set guardrails or failing to connect a customer with a person when it can’t help.
One of the best-known examples is probably DPD’s AI assistant, which in 2024 failed to track a missing parcel, then failed to connect the customer with a human representative, and ended up writing him a poem about how unhelpful it was. That’s the kind of failure that generates immediate customer complaints and public shaming on social media.
But an unhelpful AI agent is harder to spot when it provides a wrong answer with total confidence. In this scenario, nobody asks to be transferred, because they have no reason to. In February 2024, a British Columbia tribunal ordered Air Canada to pay damages after its website AI assistant told a grieving customer he could claim a bereavement discount within 90 days of booking.
The company’s actual policy, published on its website, said otherwise, but the AI agent contradicted it anyway, and the customer was later refused the refund. The tribunal found the company liable for negligent misrepresentation, rejecting its argument that the chatbot was a separate legal entity.
These failures haven’t stopped happening, as NBC Chicago reported this month. Confidently wrong answers are a common reason companies pull back AI agents handling customer communications. In the our survey, 22% of the organizations running live AI agents had rolled one back over hallucination or brand risk.
Has a deployed AI agent ever been rolled back or shut down due to a governance failure?
When asked about the most significant business impact when an AI-driven customer interaction fails, 34% of organizations chose reputational damage and lost customer trust. This is the kind of damage that might never be reversed.
Some customers will never ask for a person, because nothing in the conversation suggests they should. For the company running the AI agent, the challenge is defining when a person should take over regardless, and making sure it actually happens, whether the agent has run out of useful answers, wasn’t equipped to give a correct one, or shouldn’t be the one answering at all.
of enterprises have rolled back an AI agent over hallucination or brand risk (Sinch, 2026).
name reputational damage and lost customer trust as the biggest impact of an AI failure (Sinch, 2026).
Reading the room: sentiment analysis explained
Sometimes, messages get shorter. Sentences clip, capitals appear, the please drops off the end. A conversation that has got to that point is often one that should have been handed over already.
Sentiment detection assesses the emotional tone of each message in real time, alongside intent recognition, which detects what the customer is trying to get done. Together, they inform the agent’s next step – whether someone should step in, and who, before the customer thinks to ask, or gives up and never comes back.
By scoring emotional weight, sentiment analysis also helps ensure the most pressing customer queries move up the support queue. An AI agent can escalate, flag a case as urgent, and transfer it to the right team. This means a furious customer charged for a booking they never made goes to billing. A customer whose order is weeks late goes to the team who can trace it. Neither waits behind someone asking about opening hours. And whoever takes over needs the conversation as it stands, with full history, to avoid creating additional friction.

Sentiment is one input, not the whole decision. An irritated customer might be one reply away from an answer an AI agent can give. What happens when frustration is detected is a design choice, not the agent’s call.
Drawing the line: what an AI agent should and shouldn’t handle
An AI agent’s error rate can be reduced but never brought to zero, so what matters is where its limits are set, which questions it answers and which it passes on. That’s a company decision made before the agent goes live rather than a judgment made inside the conversation. The narrower the scope, the fewer ways a wrong answer turns into a real problem for a customer.

An agent whose scope is narrowed too far, though, is no help at all, since most of what makes it useful is the amount of context it’s allowed to work with.
Some subjects still naturally belong with a person regardless of what the AI agent knows. Ahead of Black Friday and Cyber Monday 2026, we asked 2,501 consumers across 8 countries which tasks they would trust AI with. Responses show that while overall confidence in AI agents is high, it tends to drop the closer the task gets to money and account security.
How confident you are that a brand’s AI assistant can handle the following?
Instead of scoping those out entirely, companies deploying these agents should make the handover easy rather than something the customer has to fight for.
“I don’t know” might be the smartest answer
It’s still tempting to read an AI agent admitting it can’t help as a system that’s failed, and the way customer support performance gets measured encourages this. Containment and deflection rates track how many conversations AI handles without a person stepping in, and teams are scored on keeping them high. This means the measured performance goes down when a customer reaches someone who can actually help.
The same logic applies when an AI agent gets pulled back after a failure. It looks like governance hasn’t paid off, where it’s actually the program working as it should. Across all organizations running AI agents in customer communications, 74% have done it at least once. Where it gets interesting is that among those describing their AI safeguards as fully mature, the number rises to 81%.
The cause boils down to monitoring rather than performance: fewer rollbacks may just mean poorer visibility into them. Which means the first signal telling a company its AI agent has failed is somebody complaining to customer service… or out on social media.
For a customer trying to sort out a problem, an AI agent that admits it can’t help and hands over to a human is worth more than one that pretends it can. How the two work together is what makes a difference to the customer experience, not which one the resolution came from.
Sinch builds the layer where smart conversations happen, reading intent and sentiment as each message arrives, and carrying the handover through, across messaging, voice, and email, with compliance and delivery built in. Discover how our infrastructure powers AI-driven conversations like no other can.