Evaluation guide
What to measure in a voice-agent pilot
Call count and a convincing voice are weak proxies for business value. A pilot should show that conversations produce correct outcomes and controlled next steps.
Separate conversation quality from workflow quality
Natural speech matters, but it does not prove that the agent understood the customer or completed the right next step. Review both the interaction and the resulting business state.
Measure the outcome
Define the valid dispositions before launch and sample whether they match what actually happened.
- Conversation reached a valid outcome
- Required information was captured
- Disposition and reason are correct
- The customer’s stated preference was preserved
Inspect actions and handoffs
Check whether scheduled steps, records, follow-up, and routing match the workflow rules. For handoffs, verify that the receiving person has enough context to continue without asking the customer to start again.
Review failure honestly
Track ambiguity, tool failures, unavailable context, unexpected requests, and customer frustration. The pilot should expose where rules or ownership need improvement before more volume is added.
