How Do You Test AI Features That Give a Different Answer Every Time?
You wrote a test for your new AI feature. It asserts that the model, given a customer question, returns a specific helpful answer. It passed on Monday. It failed on Tuesday—same input, different wording, still correct. By Wednesday your team had wrapped it in a retry loop that runs until green.
That’s not a test anymore. That’s superstition with a CI badge.
The core problem is that you’re applying deterministic testing instincts to a non-deterministic system. A traditional function maps one input to one output, forever. An LLM maps one input to a distribution of plausible outputs, any of which might be correct. Even with a fixed seed and temperature zero, providers do not guarantee deterministic outputs—repeated requests can still differ. The moment you assert on an exact string, you’ve guaranteed a flaky test—and flaky tests train your team to ignore red, which is far more dangerous than having no test at all.
This is a different problem from testing a mobile app, where the same tap should always produce the same screen. With AI features, variation isn’t a bug to eliminate. It’s the nature of the thing you’re testing. So you stop fighting it and change what you assert on.
Test Properties, Not Strings
The single most important shift: assert on properties of the output, not the literal output.
The model can phrase an answer a hundred ways. But a correct answer shares structural properties across all hundred phrasings. Those properties are what you test.
For a customer-support answer, the properties might be:
- Structural validity. Is it well-formed JSON? Does it match the schema you promised downstream consumers?
- Groundedness. Does it cite a source that actually exists in your knowledge base, rather than hallucinating one?
- Boundedness. Is it under your token and cost budget? An answer that’s correct but 4,000 tokens long is a production incident waiting to happen.
- Safety. Does it avoid leaking PII, giving prohibited advice, or going off-topic?
- Relevance. Does it actually address the question, measured by semantic similarity to a reference answer rather than exact match?
None of these care about exact wording. All of them fail loudly when the model genuinely misbehaves. That’s the difference between a test that protects you and a test that annoys you.
Build Eval Sets, Not Just Unit Tests
A single assertion on a single input tells you almost nothing about a probabilistic system. You need volume. This is where eval sets come in—and if you take one practice from this post, take this one.
An eval set is a curated collection of input examples paired with the properties a correct response must satisfy. Think dozens to hundreds of cases, not one. You run the whole set, score each case, and report a pass rate. You’re no longer asking “did this test pass?” You’re asking “what percentage of my golden cases satisfied their required properties, and did that percentage drop since the last change?”
A good starting eval set covers:
- Happy-path cases that represent your most common real inputs.
- Edge cases—empty inputs, very long inputs, adversarial phrasings, multilingual inputs if relevant.
- Known failure cases—the specific inputs that broke in production before. Every incident becomes a permanent eval case so the same failure can never silently return.
- Regression cases that lock in behavior you’ve explicitly decided is correct.
We typically tell teams to start with 30-50 cases that they can curate by hand and trust completely, then grow toward a few hundred as the feature matures. Quality of cases beats quantity every time. Fifty cases you understand deeply are worth more than five hundred you scraped and never reviewed.
Use Tolerance Bands, Not Hard Gates
Because the system is probabilistic, a single run might pass 47 of 50 cases and the next run might pass 48. If your CI fails the build the instant any case fails, you’re back to flaky hell.
Instead, gate on a tolerance band. Set a pass-rate threshold—say, 92% of eval cases must pass—and fail the build only when you drop below it. A single non-deterministic miss doesn’t break your pipeline; a real regression that drags the pass rate down does.
Two refinements make this robust. Run each case more than once for properties where determinism matters—if you sample a case three times and it passes twice, that’s signal about reliability, not noise to suppress. And track the trend, not just the snapshot. A pass rate that’s been 94% for a month and suddenly reads 89% is a regression worth investigating even if it’s above some absolute floor. The delta is often more meaningful than the level.
When the Judge Is Also an LLM
Some properties—tone, helpfulness, whether an answer “actually resolves” a question—can’t be checked with a regex. The emerging practice is to use a second model as a judge: you prompt an LLM to score the output against a rubric.
This works, but it introduces its own non-determinism, so treat it carefully. Three rules we follow:
Pin the judge. Use a fixed model version and a low temperature for the judge so your scoring is as stable as possible. A judge that drifts makes your whole eval set untrustworthy.
Validate the judge against humans. Periodically have a person score a sample of the same outputs and confirm the LLM judge agrees—measuring agreement against a human-labeled holdout set is how you know the judge approximates human judgment. If the judge and your humans diverge, the judge needs a better rubric before you trust it in CI.
Use the judge for graded properties, not safety-critical ones. Whether a response leaked a credit card number is a deterministic check you should never delegate to a probabilistic judge. Reserve the LLM judge for genuinely subjective qualities.
Separate the Two Sources of Failure
When an AI feature test fails, there are two very different culprits, and conflating them wastes days.
Your code changed. A prompt edit, a retrieval bug, a context-window truncation, a parsing error in how you handle the response. These are real regressions and your job is to catch them.
The model changed underneath you. Provider model updates can shift behavior overnight—OpenAI notes that its system_fingerprint changes when the backend is updated, which breaks reproducibility. An eval suite that was passing at 94% can drop because the upstream model was upgraded, not because your code is broken.
This is exactly why eval sets and version pinning matter. Pin the model version explicitly so an upstream change is a deliberate event you test for, not a surprise in production. When you do upgrade the model, run your full eval set first and treat any pass-rate drop as a blocking finding. The eval suite becomes your contract with a dependency you don’t control.
A Practical Starting Point
If you’re standing up testing for an AI feature in the next 4-8 weeks, here’s the sequence we’d run.
- Define properties. For your feature’s output, write down the 4-6 properties a correct response must have. Make as many as possible programmatically checkable.
- Build a 30-50 case eval set drawn from real and adversarial inputs, including every past failure you can remember.
- Set a tolerance band—a pass-rate threshold your CI gates on—rather than demanding 100% on every run.
- Pin your model version so upstream changes are deliberate.
- Feed incidents back in. Every production failure becomes a permanent eval case.
At Particle41, when we build LLM-powered features, the eval set ships alongside the feature as a first-class deliverable—because a non-deterministic capability without an eval harness is something you’re hoping works, not something you know works. Senior engineers paired with AI agents move fast on the implementation; the eval suite is what lets us move fast without breaking the part the business depends on.
You can’t pin the output of an AI feature. But you can absolutely pin its quality. Stop testing the string and start testing the system—and your CI will go back to meaning something.