human[it]loop

evaluation

I'll break your bot before your users do.

Synthetic users push your AI product against its hardest cases. A judge model scores the runs. Then the part nobody else sells: a human reads every transcript.

Why a human reader

An eval score is a number. The worst transcript is an insight. Automated red-team platforms will sell you the score for five figures; I sell the reading — every conversation, read by a person, written up as a field report your whole team will actually finish.

What you get

An agreed scenario set run against your bot (your API keys, your environment), judge-model scoring for the numbers, and a narrative field report — the format proven publicly on my own AI product's research pages — plus a prioritized fix list.

Scope and cost

A focused stress round runs 30–45 hours (€1,500–2,250). Continuous monthly rounds with regression tracking are available once the first round lands. For comparison: security vendors' entry audits for a simple chatbot start around $8–15k.

Write to a human

I read every inquiry myself. No form-fed CRM autoresponder on the other end.