AI Evaluations
AI Evaluations grade every conversation against quality criteria you define yourself — automatically, the moment the conversation ends.
You describe the behavior your agent must show (for example "Stay booked: each time the Guest wanted to book a stay, the agent booked it"), and Voicetta checks every call and every WhatsApp/SMS thread against it.

Why it matters for your business
- Standardize service: define the behaviors your agents must show and check them on every conversation, not a random sample.
- Work faster: no more listening to calls one by one — every conversation is graded automatically with a short explanation.
- Save money: spot failing behaviors early, before they cost bookings or repeat Guests.
- Delight every Guest: measurable conversation quality means problems are fixed based on data, not guesswork.
How an evaluation works
An evaluation is a short title plus a description of the expected behavior. After each conversation, Voicetta compares the transcript against every enabled evaluation and returns one of three results:
- Pass — the expected behavior was shown every time it was relevant.
- Fail — the situation occurred but the agent did not behave as expected.
- Not applicable — the situation never came up in this conversation.
Every result includes a one-sentence rationale explaining the verdict.
One evaluation, all your agents
Evaluations are shared across your whole workspace:
- every evaluation is visible to all agents on the AI Evaluations page,
- in the evaluation settings you choose per agent whether that evaluation is turned on or off,
- the switch on the evaluation list toggles the evaluation for the currently selected agent,
- badges on each evaluation card show which agents currently run it.
This way you define a quality standard once and decide exactly where it applies.
Voice calls
Calls are graded the moment they end. Calls without a usable transcript are marked not applicable rather than guessed.
WhatsApp and SMS threads
Text messages are grouped into conversation threads. A thread is graded as a whole 30 minutes after the last message, so the verdict reflects the full exchange instead of a single message. Each graded thread also gets a Thread Summary visible in its details.
If the Guest replies later, the thread reopens and is re-graded — a fail can flip to a pass.
Recovery follow-up (optional)
For each evaluation you can enable a recovery follow-up. When a thread fails because the Guest started a process but did not finish it (for example an abandoned booking), the agent sends one short, friendly message inviting them to complete it — in the Guest's own language.
Safety rules built in:
- at most one follow-up per thread and one per Guest per day,
- Guests who opted out never receive it,
- the message never claims anything was completed,
- WhatsApp follow-ups respect the 24-hour messaging window.
Where you see results
- AI Evaluations page: aggregate pass rates and pass/fail/not-applicable counts per evaluation, with a drill-down into individual results and CSV export.
- History: every call and thread detail includes an evaluations section with the verdict and rationale per criterion.
- Guest timelines: evaluation results appear on the Guest's conversation cards, including "Recovered" badges when a follow-up turned a fail into a pass.
Step-by-step in the app
- Open AI Evaluations from the main menu.
- Click Add Evaluation and enter a title and a description of the expected behavior.
- Choose which agents should run this evaluation.
- Optionally enable the recovery follow-up for WhatsApp/SMS.
- Save — new calls and text threads are graded automatically from now on.
Important notes
- Evaluations grade conversations that end after the evaluation was created and enabled. Existing history is not graded retroactively.
- Grading is transcript-based and never affects the call itself, saving, or billing.