AI + Assurance

Can you prove your AI agent works?

Huceptron InsightsBy the Huceptron senior partners·6 min read

A demo is a sample of one, chosen by the person selling it. If an agent is going to touch customers, money or a regulator, somebody has to be able to say how well it works and show the working.

01 What a demo proves

That the system can produce a good answer to a question someone picked. That is worth something, and it is not evidence. Evidence is what happens on the tasks nobody chose, on the day nobody prepared for, at the volume you will actually run.

02 What evidence looks like

A task bank written before the system is tested, covering the work the agent will really do, including the awkward cases. Blind scoring, so the person marking does not know which system produced which answer. A held-out set the builders never saw, because anything tuned against its own test set will pass it. Stated criteria before the run, so the bar cannot move afterwards to fit the result. And failures recorded, not just successes - an evaluation that only ever reports good news is marketing with a chart on it.

03 Five questions a board should ask

What proportion of real tasks does it complete correctly, and who measured that? What does it do when it does not know - stop, or guess confidently? Who is accountable when it is wrong, and what does the customer see? What evidence would we hand a regulator tomorrow morning? And when did we last re-run this, given the model underneath changes without asking us?

04 Why the vendor cannot answer these alone

Not because vendors are dishonest, but because the questions are about your tasks, your tolerance for error and your obligations. A supplier can tell you how the system performs in general. Only you can define what counts as correct in your business, and that definition is most of the work.

05 What good looks like

An evaluation that a sceptical outsider could repeat and get the same answer. Written method, fixed task set, blind marking, held-out cases, results kept whether or not they flatter the system. It is not exotic - it is how any other engineering claim gets tested - and it is the difference between believing an agent works and being able to show it.

We apply the same method to our own d‑Employees before we apply it to anything else. Related: you bought the AI, you still own the obligation.

AI assuranceEvaluationAI agentsEU AI ActGovernance
One week. Then you know.
The AI audit week - what applies to you, where the gaps are, and what to do first
See what the week covers →