evaluation
I'll break your bot before your users do.
Synthetic users push your AI product against its hardest cases. A judge model scores the runs. Then the part nobody else sells: a human reads every transcript.
Why a human reader
An eval score is a number. The worst transcript is an insight. Automated red-team platforms will sell you the score for five figures; I sell the reading — every conversation, read by a person, written up as a field report your whole team will actually finish.
What you get
An agreed scenario set run against your bot (your API keys, your environment), judge-model scoring for the numbers, and a narrative field report — the format proven publicly on my own AI product's research pages — plus a prioritized fix list.
Scope and cost
A focused stress round runs 30–45 hours (€1,500–2,250). Continuous monthly rounds with regression tracking are available once the first round lands. For comparison: security vendors' entry audits for a simple chatbot start around $8–15k.
Write to a human
I read every inquiry myself. No form-fed CRM autoresponder on the other end.