ai guest / customer communication

How to build automated conversation grading for calls and WhatsApp

Turn support standards into evidence-based pass/fail rules, then test the scorecards against real calls and message threads before scaling them.

By Javier Solis·October 6, 2026·4 min read
What matters here
  1. A useful scorecard defines observable behaviors and the transcript evidence that proves each one.
  2. Keep channel-specific rules for voice and WhatsApp alongside a small shared set of service standards.
  3. Review false passes and false failures with human reviewers before relying on automated grading at scale.

Automated conversation grading is only as useful as the rules behind it. “Be helpful” is a service aspiration, not a test. A scorecard needs observable behaviors, clear evidence, and a decision rule a reviewer can apply consistently to a phone transcript or a WhatsApp thread.

Voicetta says it can grade every call and every WhatsApp thread, with pass, fail, and a reason. To put that capability to work, start by writing the rubric—not by trying to grade everything your support team does. Voicetta is described as done-for-you, so bring a draft rubric to its setup conversation and agree how each behavior should be evaluated.

1. Choose the conversations and outcomes

Pick one support queue or customer task first, such as appointment changes, order questions, or new-customer inquiries. Define what a good outcome means for that task. It might mean the customer received an accurate answer, the agent gathered required details, or a case needing specialist attention was escalated.

Keep the first rubric short. Start with a few behaviors that affect accuracy, customer effort, or risk. If every desirable quality becomes a rule, the scorecard will be hard to interpret and harder to improve. Record the scope too: which conversation types are included, and which should be excluded because the rubric does not apply?

2. Write rules that can be proved from the conversation

For each behavior, write three things: what the agent must do, what transcript evidence counts, and what should cause a fail. “Show empathy” is too open to interpretation. “Acknowledge the customer’s stated problem before asking for information” is more testable. “Give the correct answer” also needs a reference point: identify the approved policy or source the answer must match.

Separate critical requirements from coaching standards. A missing required disclosure or an incorrect high-risk instruction may need to fail a conversation on its own. A less consequential behavior, such as a missed offer of additional help, may be useful for coaching without invalidating an otherwise successful interaction. State those distinctions in plain language rather than leaving them to the grader to infer.

Ask reviewers to point to a phrase or exchange as evidence. If they cannot agree what evidence proves the rule, rewrite it. Rubrics that replace manager instinct with explicit criteria are the subject of Kulissa’s discussion of automated rubric scoring; that same discipline helps make automated call transcript scoring more dependable.

3. Keep shared standards and channel-specific checks

Use a small common core across channels: accuracy, respectful treatment, and appropriate escalation, for example. Then add rules that reflect how each channel works. A phone call may require checking whether the agent confirmed a name or repeated an appointment time clearly. A WhatsApp exchange may need a rule about answering the customer’s actual question before sending another message or closing the thread.

Do not grade a message thread as if it were a single spoken turn. Read the exchange in context: a short reply may be sufficient after the customer has already provided details, while the same reply could be confusing at the start of a conversation. Likewise, do not assume a phone transcript captures everything that matters. If tone or an audio event is part of the standard, establish whether the available evidence supports grading it before making it a pass/fail rule.

4. Define the pass/fail decision

Write down how individual rules combine into an overall result. Specify which failures are automatic fails and which can be balanced against other behaviors. Avoid an unexplained overall score: a team lead should be able to see what passed, what failed, and why. Voicetta’s stated pass/fail-and-reason format can make those decisions easier to discuss, but the rubric still has to define what counts as a reason.

For each rule, include a few short examples: a clear pass, a clear fail, and a borderline case. Use examples from both voice transcripts and WhatsApp threads where the rule applies to both. This gives the people reviewing the rubric—and the automated grading setup—a shared interpretation to test against.

5. Test, calibrate, and revise

Before using results to assess performance, take a varied sample of real conversations and have experienced reviewers apply the draft rubric. Include straightforward cases, edge cases, and conversations where the customer’s request changes midway. Compare reviewer decisions with the automated grades. For each mismatch, ask whether the transcript lacks evidence, the rule is ambiguous, or the result is genuinely wrong.

Track false passes and false failures separately. A false pass can hide a service or policy problem; a false failure can teach staff to optimize for a badly written rule. Revise the rule, then check it against new examples. Keep a record of changes so a shift in results is not mistaken for a shift in agent quality. The earlier Voicetta category digest on QA metrics provides related background on automated grading measures.

Once the rubric is stable for one queue, extend it carefully. A customer support director should be able to explain each rule, show the evidence behind a result, and name the action a fail triggers—coaching, a policy review, or escalation. Automated conversation grading can broaden quality assurance beyond the calls managers have time to review. It cannot make a vague standard fair. Clear criteria, tested across phone and WhatsApp, are the work that makes scale possible.

More from Voicetta News