Back to events

Upcoming · Human in the loop

Evals in practice: how to know your agent is getting better

A hands-on workshop building an evaluation harness for a customer-support agent, from first test case to CI gate.

Register free

Teams ship agents and then have no way to tell whether last week’s prompt change helped or hurt. This workshop fixes that.

We build an eval suite live: collecting real failure cases, writing graders, setting a pass threshold, and wiring the whole thing into CI so a regression blocks the merge.

Speakers

Head of AI, Enterprise
Applied ML and evaluation

Start before the gap gets expensive.

Tell us where the work is stuck. We will tell you, in plain language, what AI can and cannot fix, and what it would take to do it properly.