Every other part of eChat is about building your agent. Evaluate is about checking your work. It runs your agent's replies against a set of real past conversations and scores each answer, so you're not left guessing whether a change to training actually made the agent better or just made it different.
The comparison point can be a real person's actual reply from your conversation history, or an earlier version of your own agent's replies. Either way, an automatic judge scores each exchange and shows you exactly where the agent's answer lined up with the reference reply and where it drifted.
Use Evaluate any time you're about to trust a change you can't easily undo - a large batch of new Q&A, a full site re-crawl, or a switch to a different AI model. It turns "I think this is better" into a number you can point to.
What Evaluate Measures
What it is
Evaluate replays a set of past conversations through your current agent, turn by turn, and scores each of its replies against a reference answer - either what a human on your team actually said at the time, or what an earlier version of the agent said. It gives you a quality signal that isn't just a hunch.
What it's best for
- Confirming a training change actually improved answers, instead of assuming it did
- Comparing your agent's current replies against your best team members' replies
- Catching a regression before your visitors do
How it works
Evaluate takes each exchange in your evaluation set, sends the visitor's message to your live agent, and compares the reply it generates against the reference reply for that same message. An automatic judge scores how closely the two match in substance - not word for word, but whether the agent covered the same information and got it right. You get a score for every exchange, plus an overall picture of how the run went.
Building an Evaluation Set
What it is
An evaluation set is the batch of past conversations Evaluate replays and scores against. You need one before you can run anything.
How it works
You build an evaluation set the same way you'd bring in past conversations for Training > Conversations: either import directly from your eWebinar chat archive, or upload a conversation export file. Each conversation in the set becomes a series of scored exchanges - the visitor's question and the reply that was actually given, which the judge treats as the reference answer to beat. Conversations used for evaluation are only a yardstick; they don't train your agent.
If a set can't be built - for example because your filters were too strict and no conversation had enough reference replies - Evaluate shows a clear failure reason instead of spinning forever on an empty panel. Use that message to loosen the filters or pick a different source of conversations, then try creating the set again.
Best practices
- Pull in conversations that represent your real mix of questions - pricing, setup, troubleshooting - not just the easy ones
- Favor conversations handled by your strongest team members if you're using human replies as the reference; a weak reference makes every score look better than it should
- Refresh your evaluation set occasionally as your product and common questions change, so you're not scoring against conversations from a year ago
Reading Your Results
What it is
Once a run finishes, Evaluate shows you a score for every exchange in the set, not just an overall grade.
How it works
Each exchange shows the visitor's original message, the reference reply, and what your agent said instead, side by side with a score. Where the agent's answer diverges from the reference, you can see exactly why the judge marked it down - a missed detail, an incorrect claim, or an answer that was technically fine but skipped what the reference reply covered. This is where Evaluate is more useful than a single pass/fail number: it points you straight at the specific gap, so you know whether the fix is a missing Q&A pair, an outdated web page in training, or an instruction that needs tightening.
What an Evaluation Costs
What it is
Running an evaluation uses credits from the same account balance as your agent's live replies.
How it works
Each run is charged at what it actually cost to generate your agent's replies and judge them for the test cases that succeeded, plus 10%, with a minimum of 1 credit per run. Test cases that fail aren't charged. Evaluation charges appear in your usage ledger and count toward auto-recharge on paid plans. See Plans, Credits & Usage.
Best practices
Keep evaluation sets focused - a few dozen representative conversations usually tell you as much as hundreds, for a fraction of the credits. Comparing models costs more with pricier models, since the agent's replies are generated with the model you're testing.
When to Run an Evaluation
What it's best for
- After importing a large batch of new Q&A pairs, to confirm they actually raised answer quality
- After a big website re-crawl, since new or changed pages can shift how the agent answers existing questions
- Before switching your agent to a cheaper or faster AI model, to confirm quality holds up before you commit to the switch
Best practices
Don't treat a training change as done just because it's live. Run Evaluate right after, and again a week or two later once you've collected fresh real conversations, so you're checking against current behavior rather than a stale set. If you're chasing lower credits costs by moving to a lighter model, run the same evaluation set on both models before switching for good - a small dip in score might be an acceptable trade for cost, but you want that to be your decision, not a surprise your visitors notice first.
Evaluate works best alongside Chat Logs, which shows you real visitor feedback as it comes in, and Q&A training, which is usually the fastest way to close a gap Evaluate turns up.