Conversational assessment is where a learner talks to an AI avatar, in a scenario with stakes, and is scored on the judgement they show rather than the facts they can recall. It is genuinely good for high-stakes conversations that require empathy, timing and decision-making, and it falls down where the scenario has no rubric or the model is left to grade on its own. The difference between a gimmick and an assessment people trust comes down to three things: a defined scenario spine, evidence-linked scoring, and a human who reviews the result.
This is our most advanced work, and there is very little honest writing about it. So here is what we have learned about where it earns its place and where it does not.
What is conversational assessment, and how is it different from a chatbot?
A chatbot answers questions. A conversational assessment puts the learner inside a scene with a role, stakes and consequences, and grades what they do. A chatbot answers questions. A simulation creates a scene with a role, stakes, and consequences. The avatar behaves like a person in context, and the system measures the learner's actions against a rubric rather than simply providing information.
That is the whole distinction, and it matters. The learner is not being tested on whether they remember the policy. The learner is not memorizing policy. They are rehearsing judgment, communication, and timing, with feedback that is specific enough to improve next attempt.
Recall you can test with a multiple-choice quiz. Judgement you can only test by watching someone handle a situation. An avatar that plays a difficult customer, an anxious patient or a resistant witness lets you do that at scale, without booking a room full of actors.
What is it genuinely good for?
It is good for skills where human interaction is central and the cost of getting it wrong is high. That is a narrow definition on purpose. Effective use cases begin by identifying skills that require judgment, empathy or adaptive communication. The avatar should serve those skills, not define them.
The strongest cases we see are the ones where practice would otherwise be scarce, awkward or expensive. Contact-centre and customer-service work is one. Contact center and customer service agents manage sensitive conversations that require clarity, confidence and professionalism. Simulated calls create space for agents to practice without consequence, exposing trainees to high-stakes situations in a safe, repeatable environment. This boosts their readiness and resilience before they ever plug in a headset for a real call.
The same logic holds in fields where empathy shapes the outcome. In fields where empathy and communication directly affect outcomes, avatars are being used to simulate patient or client interactions. These environments allow learners to practice complex, high-stakes conversations that improve preparedness, empathy and effectiveness, leading to better outcomes for individuals in need.
For our audience, translate that into a safety conversation, a near-miss report, a supervisor challenging an unsafe act, or a first-line manager handling a difficult return-to-work discussion. These are exactly the moments where a written procedure exists but the behaviour under pressure is what actually matters.
There is also a practical assessment advantage that live role-play cannot match. In a trial assessing communication skills, researchers found the avatar system produced a full record of every interaction. Unlike live role-plays, which rely heavily on assessor observation and note-taking, the system produced full transcripts and AI-generated summaries of each interaction. This allowed assessors to review conversations in detail, revisit specific moments and evaluate performance against criteria more consistently.
That consistency is worth as much as the scale. Two assessors watching the same live role-play will remember different things. A transcript does not.
Where does it fall down?
It falls down when you trust the model to grade on its own. This is the single most important thing to understand before you buy or build one.
The research on using one AI model to score another is blunt about the limitation. Model-based evaluation relies on another model to assess the target model's performance. However, this approach is only as reliable as the evaluator model itself, which may have inherent limitations, and could result in potentially misleading evaluations.
We have seen how that goes wrong in practice, and so have others. In the Singapore communication-assessment trial, candidates built rapport by referencing local food, a culturally specific and skilful move. It was common for candidates to attempt to build rapport by referencing local dishes or asking about favourite hawker dishes, a conversational strategy typical in Singapore, where food is an important connector across communities. While human assessors viewed these efforts as positive and culturally sensitive, the AI often did not identify them as rapport-building strategies.
The humans saw good judgement. The model missed it. That is not a bug you patch once. It is a structural property of the tool, and the honest design response is to keep a person in the loop. In that same trial, the AI ran a first pass against the rubric, but that was the ceiling of its authority. The GenAI system also conducted a first-round evaluation using the assessment rubric. However, this was never intended to replace human judgement.
The second failure mode is more basic: the character does not hold. Benchmarks that test how well models stay in role show it is genuinely uneven. One evaluation scored a leading model strongly on decision-making and moral alignment but found significant weakness in In-Character Consistency for another. If your avatar breaks character mid-scenario, or leaks knowledge the character should not have, the assessment is worthless. It stops being a test of the learner and becomes a test of the model.
The third failure mode is the one that turns the whole thing into a gimmick: realism without design. Despite their potential, AI avatars are not a universal solution and, in some cases, they may undermine learning outcomes if they lack realism or are designed poorly. Avatars that lack rigorous, sound learning design could derail the learning experience. A convincing digital human with no clear objective, no rubric and no reflection is just a more expensive way to do nothing.
What makes an assessment people actually trust?
Three things, and none of them is the avatar itself.
First, a scenario spine. The avatar is allowed to improvise, but the scene is not free-form. Use beat based structure, controlled variation, and role boundaries. The avatar can improvise inside a defined scene, but the scenario spine keeps the experience comparable for assessment. Without that spine, two learners get two different tests and you cannot compare their scores.
Second, evidence-linked scoring and a clear decision about what you are grading. There is a genuine fork here worth naming. This offers a choice: do you score the conversation itself, or do you use the conversation to provide context for a set of related, scored test questions. The second option is often the more defensible starting point for a regulated programme, because the conversation itself is not what gets scored. The conversation is context that is used to answer any type of related item, including multiple-choice or short-answer questions. In this case, it's the linked questions that are scored, as opposed to the conversation. It keeps the realism where it belongs and keeps the scoring somewhere you can already defend.
Third, an audit trail and a human review workflow. For anything that touches compliance or certification, this is not optional. Can these simulations be used for compliance certification? Yes, if the rubric is explicit, the scoring is evidence linked, and the system supports audit trails, confidence thresholds, and review workflows. Regulated programs should also document governance and content boundaries.
“A convincing digital human with no rubric behind it is a more expensive way to do nothing.”
When is a conversational avatar the wrong answer?
When you are testing recall, use a quiz. When the skill is procedural and has one correct sequence, use a checklist or a hands-on assessment. When there is no rubric because nobody can agree what good looks like, fix that first, because an avatar will not fix it for you.
And if the honest reason for wanting an avatar is that it looks impressive in a board demo, stop. This audience has been sold impressive-looking things for years. The version worth building is the one where a real assessor would sign their name to the result, and where you can show them the transcript when they ask why. If you cannot get to that, the conversation is a rehearsal tool, which is valuable, but it is not an assessment, and you should not call it one.
That honesty is the whole game. Conversational assessment is powerful in a narrow band of high-stakes, human-facing skills, and hollow everywhere else. Knowing which is which is the work.
