What to Ask in an AI Support Demo (So You Don't Get Fooled)
Great demos hide weak products. These questions reveal whether the AI is real and safe.
A great demo is not the same as a great product. Vendors rehearse the happy path, pick the easy questions, and lean on a slide deck for the hard parts. The way to see through it is to bring your own questions — ones that probe what happens when things go wrong, not when they go right.
Here are the AI support demo questions worth asking, grouped by what they expose. Ask them live, in the product, and watch whether the answer is a demo or a deflection.
Does the AI actually know your stuff, or is it guessing?
The single biggest failure mode of an AI support agent is confidently making things up. Your first job is to find out where its answers come from.
- "Where did that answer come from — show me the source." A real product cites the doc, the order record, or the policy it used. If it can't, it's improvising.
- "Ask it something we've never documented." A safe agent says it doesn't know and hands off. A weak one invents a plausible-sounding answer.
- "Ask it about an order that doesn't exist." It should refuse cleanly, not fabricate a status.
- "How does it stay current when our policy changes?" You want to hear that it reads from your live systems, not a snapshot someone pasted in months ago.
If the AI can't show you the receipt for its own answer in the room, assume it's guessing in production too.
What can it actually do — and what stops it from doing the wrong thing?
Answering is easy. Acting is where the risk lives. An agent that can issue a refund can also issue the wrong refund, twice. Push on the guardrails.
- "What actions can it take on its own, and which need a person?" There should be a clear, configurable line — not "it figures it out."
- "Show me it trying something it's not allowed to do." A safe product blocks the action and explains why. If everything is permitted, nothing is actually controlled.
- "What happens if it fails halfway through an action?" You want fail-closed behavior — it stops and flags, rather than leaving a customer half-refunded.
- "Can it accidentally do the same thing twice?" Ask specifically about double refunds and duplicate orders. The honest answer involves protection against repeats, not "that won't happen."
If a vendor can't make their own AI fail safely on demand, they've never tried — which means you'll be the one who finds the edge in production.
Can you audit what it did after the fact?
When a customer disputes something the AI did, "the model decided" is not an answer you can give. You need a record.
Ask: "After the AI takes an action, what can I see?" The strong answer is a receipt — what the AI saw, what it decided, what it was allowed to do, and what it actually did, attached to the conversation. The weak answer is a log of API calls you'd have to reverse-engineer, or nothing at all.
Then ask: "Can a non-engineer read it?" The whole point of a receipt is that your CX lead, not just an engineer, can answer "why did this happen" without filing a ticket.
What happens at the edge — handoff and escalation?
The AI will hit its limit. What matters is whether it knows when, and whether the person who takes over starts from zero.
| What to ask | Strong answer | Weak answer |
|---|---|---|
| "When does it hand off to a person?" | Clear triggers — low confidence, sensitive topic, an explicit ask | "It almost always resolves" |
| "What does the rep see on handoff?" | Full context and a summary, no re-asking the customer | A fresh ticket the rep has to read cold |
| "Can a person take over mid-conversation?" | Yes, smoothly, with the thread intact | Only at the end, or not at all |
| "Does it know when it's unsure?" | It defers and says so | It answers everything with the same confidence |
An agent that never escalates isn't confident — it's reckless. The good demos are the ones where the AI chooses to bring in a person.
Are the metrics honest?
Finally, interrogate the numbers, because this is where the best slides hide the worst news.
Ask how they measure deflection. A vendor who counts every conversation the customer didn't escalate as "deflected" is counting people who gave up. Honest deflection means the customer's problem was actually solved and satisfaction held steady. Ask to see deflection and satisfaction on the same chart. If they only show one, ask why.
Then ask about quality: "Do you score every AI conversation, or a sample?" Sampling AI quality at 2% misses the same patterns it misses for humans. And ask "what's your worst category?" A confident team will tell you.
The questions above are the ones we'd want a buyer to ask us. BearScope is built so the answers hold up: every agent action is checked before it runs and leaves a receipt you can audit, the AI is grounded in your real systems, and it hands off cleanly when it should. See how the product works, or book a walkthrough and bring this list with you.
See it on your own conversations.
Bring your busiest day. We'll score every conversation in it.
Book a walkthrough →