InsightsHow to evaluate an LLM system before you trust it

How to evaluate an LLM system before you trust it

A handful of good answers in a demo isn’t evidence. Here’s how to build the evidence before real people depend on the system.

Large language models are persuasive by design. They answer fluently, sound confident and handle the examples you try in a demo. That’s exactly why evaluating them takes discipline: their failures are rarer, quieter and easy to miss until someone acts on one.

Evaluation isn’t a test you pass once. It’s a set of habits that runs from the first prototype through every change after launch.

Define the job, and how it can fail

Start by writing down what the system is for, in plain language: who asks it things, what they ask and what a good answer lets them do next. Then list the ways it can go wrong. Some failures are merely annoying, like a clumsy summary. Others are expensive: a confident answer that cites a policy that doesn’t exist, data shown to someone who shouldn’t see it, or an action nobody approved.

Be specific about scope as well. A support assistant that answers billing questions shouldn’t attempt legal advice, and deciding what the system should decline is part of defining the job.

Rank those failures by what they would cost you. The ranking decides where to spend evaluation effort, and which mistakes you can tolerate at what rate.

Build an evaluation set from real cases

Public benchmarks describe a model in general, not your task. The most useful evaluation set is made of real inputs from your domain: questions from actual support tickets, the documents your teams really use, the edge cases your experts remember.

  • Cover the common cases, the hard cases and the cases that must never go wrong.
  • For each case, write down the expected answer, or what a good answer must include and must avoid.
  • Keep part of the set aside and never tune against it, so you can tell real improvement from memorizing the test.

Start small. A few dozen well-chosen cases beat thousands that nobody has read. Add a new case every time you find a new failure.

Combine automated checks with human review

Some qualities can be checked automatically: whether the output follows the required format, whether cited sources exist and actually support the claim, whether sensitive fields appear where they shouldn’t. Using a second model to grade answers can help at scale, but treat its scores as a reason to look closer, not a verdict, and check them against human judgment regularly.

For the failures that matter most, keep people in the loop. Domain experts reviewing a sample of real outputs every week will catch problems no metric was designed to see.

Test every change like a release

Prompts, retrieval settings, model versions and source documents all change behavior, sometimes in places you didn’t touch. Run the full evaluation set on every change and compare the results with the last version you trusted. A fix for one case can quietly break others, and only regression tests catch that before your users do.

Keep the history of results, too. When quality shifts, you want to see exactly which change caused it and roll back in minutes, rather than debate it for days.

Keep evaluating in production

Real users will ask things your evaluation set never imagined. Log inputs and outputs with the privacy controls your data requires, review a sample on a regular schedule, and watch signals like user corrections, escalations and abandoned sessions. The interesting failures go back into the evaluation set.

Set a review rhythm and stick to it. A weekly look at a small sample, owned by a named person, beats an annual audit that nobody remembers to run.

Put guardrails where mistakes are expensive

  • Give the system the least access it needs: read-only unless an action is truly required.
  • Require confirmation before anything that can’t be undone.
  • Ground answers in approved sources, and let the system say “I don’t know.”
  • Route low-confidence or high-risk cases to a person.

Guardrails aren’t a sign of distrust in the model. They’re how you decide, deliberately, which mistakes the system is allowed to make on its own and which ones need a person.

Trust in an AI system should be earned the way trust in any system is: with evidence, collected continuously.

Done well, evaluation doesn’t slow a team down. It’s what lets you change prompts, switch models and ship improvements with confidence instead of crossed fingers.

If you’re about to put an LLM system in front of customers or employees and want a second opinion on how it’s evaluated, let’s talk.

Keep reading

4 min read

The first 30 days of an AI project

The first month decides whether an AI project becomes a working system or a slide. Here’s how we spend it.

4 min read

Why AI pilots stall before production

When a pilot stalls, the model is rarely the problem. The trouble is everything the demo didn’t have to handle.

Bring us the project that matters most.

Tell us what you’re building and what’s in the way. You’ll hear back within 24 hours.