Customer service

AI Support Agent Evaluation Checklist: What to Test Before You Buy

Jedrzej Meder4 min read

AI Support Agent Evaluation Checklist: What to Test Before You Buy

AI Support Agent Evaluation Checklist: What to Test Before You Buy

Every AI support agent demos well. The vendor picks the questions, the knowledge base is clean, and the answers land. Then you connect it to your own content, your own customers ask their own strange questions, and the picture changes.

This checklist is what to test yourself, in your own account, before you commit. Work through it in a trial and you will know within a week whether a product fits your support operation.

1. Knowledge ingestion

Everything downstream depends on what the agent can read.

  • Can it crawl your help center and keep it in sync when articles change?
  • Can it take files, not just URLs? PDFs and internal docs matter for policy answers.
  • Can you add plain-text answers directly for things that live nowhere else?
  • What happens when two sources disagree? A good product tells you. A weak one silently picks one.
  • How long does a content change take to show up in answers?

Test it: update one help article with a deliberate change, then ask the agent about it an hour later.

2. Answer quality on your real questions

Do not test with questions you invent. Pull the last 100 real conversations from your inbox and replay a representative sample.

Score each answer as one of four outcomes:

OutcomeWhat it means
Correct and completeThe customer would be done
Correct but thinRight, but they would ask a follow-up
DeflectedHanded off when it should not have been
WrongThe dangerous one, count these separately

A product with a slightly lower resolution rate and zero wrong answers is the better buy. Wrong answers cost trust, refunds, and sometimes compliance exposure.

3. Guardrails and control

Ask what the agent will refuse to do, and confirm it.

  • Can you block topics outright, such as legal, medical, or pricing negotiation?
  • Can you force it to answer only from your sources rather than general model knowledge?
  • Can you review and approve what it says before it goes live?
  • Does it invent policy when your content is silent?

Test it: ask about a policy you have deliberately not documented. The right answer is a graceful handoff, not a plausible invention.

4. Handoff to your team

The transfer to a human is part of the product, not an afterthought.

  • Does the whole transcript travel with the conversation?
  • Can you define your own escalation triggers?
  • What happens outside business hours?
  • Does the customer stay in the same window, or start over somewhere else?

5. Channels and coverage

List the channels your customers actually use, then check each one individually. Support for a channel on a pricing page is not the same as feature parity on that channel. Ask specifically whether knowledge, escalation rules, and reporting behave the same on email as they do in chat.

Same question for languages. Answering in a language and being genuinely accurate in it are different things, and the gap usually shows in your product-specific vocabulary.

6. Integrations with your real stack

An agent that can only talk cannot resolve much. The useful ones can look something up and act.

  • Order status, subscription state, account details
  • Your helpdesk and ticketing system
  • Your ecommerce platform, if you run one
  • Whatever internal system holds the answer customers ask for most

Test it: ask a question that requires a live lookup, not a knowledge-base answer.

7. Reporting you can act on

A dashboard that only shows conversation volume will not help you improve anything. Look for:

  • Resolution and escalation rates over time
  • The questions the agent failed on, grouped by topic
  • Content gaps surfaced automatically
  • Exportable data, so your analysts are not locked into someone else's charts

The failed-question list is the single most valuable screen in any of these products. If it is missing, improvement becomes guesswork.

8. Cost model

Understand exactly what triggers a charge before you sign.

  • Per conversation, per resolution, per seat, or a blend?
  • What counts as a billable conversation? A single greeting?
  • What happens when you exceed your plan mid-month?
  • Which capabilities sit behind a higher tier?

Model your actual monthly volume, not the volume on the pricing page example. Then model it at twice that, because seasonal spikes are when the bill surprises people.

9. Security and data handling

  • Where is conversation data stored, and for how long?
  • Is your data used to train shared models? Get the answer in writing.
  • What certifications and regional hosting options exist?
  • How do you delete a customer's data on request?

10. Time to value

Finally, measure how long the trial itself took you. If you needed a solutions engineer to get a first useful answer, that cost does not disappear after purchase, it repeats every time your team changes something.

Scoring the decision

Rank the ten sections by what your operation actually needs, then score each product. Most teams find the decision turns on three of them: answer quality on real questions, the failed-question reporting, and the cost model at real volume.

Run the trial with your own data, your own questions, and your own team doing the setup. A product that survives that week will survive production.