October 1, 2026

Live Chat Quality Assurance: A Scorecard for Reviewing Chats

Live chat quality assurance scorecard for reviewing support conversations

Live chat quality assurance means regularly reading a sample of real conversations and scoring them against a short, agreed checklist: was the answer correct, was the problem solved, was the tone right and did the agent follow the process. Done well, it takes an hour or two a week and shows you exactly what to fix in training, content and tools. Done badly, it becomes a box-ticking exercise that agents resent. This guide gives you a scorecard you can adapt, a sampling method and a way to turn scores into coaching, including for chats answered by AI.

What Live Chat Quality Assurance Is (and Is Not)

Metrics such as first response time, CSAT and resolution rate tell you how much and how fast. They do not tell you whether a specific answer was right. A chat can close quickly with a happy rating and still contain a wrong refund promise. Quality assurance, usually shortened to QA, fills that gap by looking at the content of conversations.

QA is not a performance review in disguise. The goal is to find patterns: a policy nobody understands, a missing help page, a canned reply that sounds cold. Individual scores matter mainly as a starting point for coaching.

Why Small Teams Need QA Too

Large contact centres have dedicated QA analysts. A team of two or three does not, and that is exactly why a light version helps. In a small team:

  • one wrong answer is repeated by everyone who copies it;
  • new people learn from whatever they see in old chats;
  • AI answers go unread because “the bot handles it”.

An hour a week of structured reading catches these problems before customers do. It also gives new agents a clear picture of what “good” looks like, which pairs well with a proper 30-day live chat agent training plan.

Building a Live Chat Quality Assurance Scorecard

A scorecard is a short list of criteria, each with a weight. Keep it to five or six categories. More than that and reviewers disagree, reviews take too long and agents cannot remember what matters.

Category Weight What the reviewer checks
Accuracy 30% Every fact is correct and matches current policy; no invented dates, prices or promises
Resolution 25% The customer’s actual question was answered or clearly handed on, with a next step
Clarity and tone 15% Plain language, right formality, empathy where needed, no jargon
Process 15% Correct tagging, priority, ticket conversion, internal notes and handoff
Efficiency 10% No unnecessary questions, no repeated information, sensible use of saved replies
Privacy and safety 5% No card numbers or passwords requested; identity checked before sharing account details

Adjust the weights to your business. A clinic might give privacy more weight; a shop might give resolution more.

Use simple scales

For each category, use a three-point scale: meets, partly meets, does not meet. Five- or ten-point scales look precise but produce endless debate about whether a chat is a 7 or an 8.

Define auto-fail items

Some mistakes should fail a chat regardless of the rest: sharing another customer’s data, asking for full card details, promising a refund the policy does not allow, or being rude. List them explicitly so there is no discussion.

Write examples for each category

For every category, write one example of “meets” and one of “does not meet” from your own chats. That single page is the most useful thing you can give a new reviewer.

How to Pick Which Chats to Review

You cannot read everything. A mix of random and targeted samples gives the fairest picture.

Random sample

Pick a fixed number of chats per agent per week at random, for example five. Random selection shows the typical experience and avoids reviewing only the problems.

Targeted samples

  • Low ratings. Every chat rated poorly deserves a look.
  • Long conversations. Very long chats often hide unclear policy or missing information.
  • Reopened tickets. If a customer came back about the same issue, the first answer probably did not solve it.
  • Handoffs from AI. These show what the assistant could not answer and how smoothly people took over.
  • New topics. After a price change or product launch, review chats that mention it.

Keep the random and targeted scores separate in your notes. Targeted samples are supposed to look worse, and mixing them makes averages meaningless.

Reviewing AI Answers as Well as People

If an AI assistant answers part of your chats, include it in QA. It is effectively a team member that talks to more customers than anyone else. The questions are slightly different:

  • Did the answer come from your own content, and was that content current?
  • Did it stay within your instructions, for example no delivery promises?
  • Did it hand over at the right moment, neither too early nor too late?
  • Did the visitor get what they needed, or did they leave after the answer?

When an AI answer is wrong, the fix is usually in the content, not the bot. Update the page it relied on. Our guide to auditing AI chatbot accuracy covers this in more detail.

Running Calibration Sessions

Two reviewers reading the same chat often give different scores. Calibration fixes this. Once a month, everyone who reviews chats scores the same three to five conversations separately, then compares and discusses the differences.

The discussion is the point. It forces the team to agree on what “clear” or “resolved” means in practice, and those agreements should be added to the scorecard examples. After a few sessions, scores become consistent enough to compare over time.

Calibration is also a good moment to check tone. The Nielsen Norman Group’s four dimensions of tone of voice are a practical way to describe the tone you want, and the UK government’s guidance on tone of voice is a solid reference for plain, respectful writing.

Turning Scores Into Coaching

A score without a conversation changes nothing. For each agent, pick one strength and one thing to improve per week, based on real examples.

  1. Share the chat, not only the score. Show the exact message and what a better version would look like.
  2. Ask first. “How would you answer this now?” often gets a better answer than the reviewer’s suggestion.
  3. Look for system causes. If three agents make the same mistake, the problem is the policy, the help page or the saved reply.
  4. Track trends, not single chats. A category that stays low for a month is worth attention; one bad chat usually is not.
  5. Celebrate good chats. Share excellent examples with the whole team. They teach more than any guideline.

A Worked Example: Scoring One Chat

Here is how a reviewer might score a typical conversation. A visitor asks whether a jacket bought three weeks ago can still be returned. The agent replies within a minute, apologises for the trouble, says “yes, just send it back”, and closes the chat.

  • Accuracy: does not meet. The shop’s policy allows 14 days for returns. The agent promised something the policy does not allow. This is an auto-fail item.
  • Resolution: partly meets. The visitor got an answer, but not the steps: no return address, no form, no mention of who pays for shipping.
  • Clarity and tone: meets. Friendly and short.
  • Process: does not meet. An exception to policy should have been handed to the team lead and noted internally.
  • Efficiency: meets. Quick and to the point.
  • Privacy: meets. No personal data requested.

The coaching point is not “be more careful”. It is concrete: check the returns page before answering timing questions, and hand exceptions to the lead. The reviewer also checks why the agent got it wrong. If the returns page buries the 14-day limit in a long paragraph, the page needs fixing as much as the agent needs coaching.

Extending QA to Tickets and E-mail

The same scorecard works for tickets and e-mail replies with two small changes. Give efficiency less weight, because written replies are expected to be complete rather than instant. Add a check for the subject line and the ticket status: was the ticket marked solved only after the customer confirmed, and was the priority correct? If your team works both chat and tickets in one inbox, review a mix of both each week so neither channel drifts.

Revisit the scorecard itself every quarter. Remove criteria that everyone always meets and add the new mistakes you keep seeing.

Running QA in Talkmio

You do not need a separate QA tool to start. Talkmio has enough built in for a small team:

  • Inbox filters such as Mio replied, Closed and Unanswered make it easy to pull samples of AI answers and finished chats.
  • Internal notes are visible only to your team, so reviewers can leave comments directly on a conversation.
  • Ratings from visitors help you find chats for targeted review.
  • Try Mio shows which knowledge pieces an answer used, which helps you trace a wrong AI answer to the page behind it.
  • Reports show the share answered by Mio alone, hand-offs, first-reply time, ratings and team performance. On the Ultimate plan and above you can export them to CSV and keep QA scores next to them in a spreadsheet.

Keep the scorecard itself in a shared spreadsheet: date, conversation link, reviewer, agent, category scores and one comment. That is enough for a year of trend data.

Common Live Chat QA Mistakes

  • Scoring only speed. Fast wrong answers are worse than slightly slower correct ones.
  • Too many criteria. A 25-line checklist gets filled in without thought.
  • Reviewing only complaints. Agents feel judged on the hardest chats alone.
  • Ignoring AI answers. The assistant may be the busiest “agent” you have.
  • No follow-up. If nothing changes after reviews, people stop taking them seriously.

Pair QA with the numbers you already track, such as CSAT for live chat. When QA scores rise and CSAT follows, you know the scorecard measures what customers care about.

Frequently Asked Questions

What is live chat quality assurance?

Live chat quality assurance is the regular review of a sample of real chat conversations against a short scorecard. Reviewers check accuracy, resolution, tone, process, efficiency and privacy. The results show where training, help content or policies need to change, which speed and volume metrics alone cannot show.

How many chats should we review each week?

For a small team, five random chats per agent per week plus targeted reviews of low-rated, reopened and handed-over chats is a good start. That usually takes one to two hours. Consistency matters more than volume: a small sample reviewed every week beats a large one reviewed once a quarter.

What should a live chat QA scorecard include?

Five or six weighted categories: accuracy, resolution, clarity and tone, process, efficiency, and privacy and safety. Use a simple three-point scale for each, define auto-fail items such as asking for card numbers, and add one real example of good and poor performance for every category.

Should AI chatbot answers be part of quality assurance?

Yes. An AI assistant often answers more chats than any person, so its mistakes spread quickly. Check whether answers came from current content, stayed within your instructions and handed over at the right time. When an answer is wrong, fix the page it relied on rather than adding rules to the bot.

Who should do the reviews in a small team?

In a team of two to five, the owner or team lead usually reviews, and agents can review each other’s chats once the scorecard is clear. Monthly calibration, where everyone scores the same chats and compares, keeps peer reviews consistent and fair and helps everyone understand the standard.

How do QA scores relate to CSAT?

CSAT shows how customers felt; QA shows whether the answer was correct and followed the process. They usually move together, but not always. A customer may rate a friendly wrong answer highly. Track both, and investigate when they diverge, because that is where hidden problems appear.

The Bottom Line

Live chat quality assurance does not need a big team or special software. A six-category scorecard, a weekly sample of random and targeted chats, monthly calibration and one coaching point per agent will improve answers faster than any new metric. Include your AI assistant in the review and fix content, not only people. Talkmio gives you the inbox filters, internal notes, ratings and reports to start this week; try it free at app.talkmio.com.


Try Talkmio on your site

Free plan, no card required.

Start free