Back to events
Upcoming · Human in the loop
Evals in practice: how to know your agent is getting better
A hands-on workshop building an evaluation harness for a customer-support agent, from first test case to CI gate.
Teams ship agents and then have no way to tell whether last week’s prompt change helped or hurt. This workshop fixes that.
We build an eval suite live: collecting real failure cases, writing graders, setting a pass threshold, and wiring the whole thing into CI so a regression blocks the merge.
Speakers
Head of AI, Enterprise
Applied ML and evaluation
Start before the gap gets expensive.
Tell us where the work is stuck. We will tell you, in plain language, what AI can and cannot fix, and what it would take to do it properly.