How Do You Know Your AI Got Worse? Build the Eval Set First
I opened a pull request that made my RAG system measurably worse and was completely certain it was an improvement. The build failed instead. What a golden evaluation pair actually contains, the four metrics that localise a failure, why 21% of my questions have no answer, how to pick a threshold from the noise floor, and the three of my own PRs it blocked.
I opened a pull request that made my RAG system meaningfully worse, and I was completely certain it was an improvement. Smaller chunks, tighter passages, less noise in the prompt — every instinct I had said yes. I spot-checked six questions, they looked fine, and I would have merged it on a Tuesday afternoon without a second thought.
The build failed instead. Context recall had gone from 0.89 to 0.71, because a chunk small enough to be "precise" was too small to contain a whole answer, and half my corpus answers span two paragraphs. Six spot-checks did not find that. A hundred and eighty did.
This post is about the least glamorous artefact in an AI system and the one I would rebuild first if I lost everything: the evaluation set, and the CI gate that makes it mean something.
Why intuition fails here specifically
Every AI change feels like an improvement
In ordinary backend work, a regression announces itself. The test goes red, the type checker complains, the endpoint 500s. You do not need discipline to notice — you need discipline to fix it.
AI systems have the opposite shape. A regression here produces a fluent, well-structured, confident answer that happens to be worse. Nothing throws. The output still looks exactly like output. And because generation is stochastic, running the same question twice gives you two different answers, so your instinct to "just check it again" produces noise rather than signal.
So spot-checking does not scale, and it does not even work at small scale, for a reason worth stating plainly: you pick the questions you already thought about. Those are the questions your last change was designed to fix.
The artefact
What a golden pair actually contains
Not a question and an answer. Four things:
{
"id": "elo-014",
"question": "How do I stop a global scope applying to one query?",
"ground_truth": "Call withoutGlobalScope(ScopeClass::class) on the query,
or withoutGlobalScopes() to remove all of them.",
"expected_sources": ["eloquent/global-scopes#removing"],
"answerable": true,
"note": "phrased as intent, not as the API name — tests semantic retrieval"
}
{
"id": "neg-007",
"question": "What is the default connection pool size in Filament?",
"ground_truth": null,
"expected_sources": [],
"answerable": false,
"note": "sounds plausible, has no answer. Filament has no pool."
}
expected_sources is what separates an eval set from a quiz. It lets you ask did retrieval find the right passage separately from did the model use it well — and those two failures need completely different fixes. Without it, every regression looks the same and you tune the prompt because the prompt is the thing you can see.
Measurement
Four metrics, because one number cannot answer four questions
A single "quality score" is a comfortable lie. When it drops you know something is wrong and nothing about where. These four localise the failure:
| Metric | The question it answers | Fails when | My gate |
|---|---|---|---|
| Context recall | Did retrieval find the right passage at all? | Chunking, embeddings, or the query itself | ≥ 0.80 |
| Context precision | Is the right passage near the top? | The reranker, or the fusion step | ≥ 0.75 |
| Faithfulness | Are the claims supported by what was retrieved? | The prompt, or the model filling gaps | ≥ 0.85 |
| Refusal correctness | Does it decline when there is no answer? | The citation gate, or its absence | ≥ 0.85 |
The pairing matters more than the individual numbers. High faithfulness with low recall means the system is honestly summarising the wrong passage. High recall with low faithfulness means it found the answer and then embroidered it. Those look identical on a single score and need opposite fixes.
Two blocked pull requests, two different metrics, two entirely unrelated causes. A single score would have told me "it got worse" twice and pointed at nothing.
The part that changed the system
The 21% of questions with no answer
Thirty-eight of my 180 pairs are unanswerable on purpose. They sound completely reasonable, they use the right vocabulary, and the corpus simply does not cover them. Writing them felt like padding. They turned out to be the most valuable part of the set by a wide margin.
Every metric looked healthy. Faithfulness 0.88, recall 0.89.
The system invented a confident answer for 76% of questions it had no business answering — and nothing in my dashboard could see it, because a hallucinated answer to a question with no ground truth is not a wrong answer. It is an unmeasured one.
Refusal correctness appeared as a number: 24%.
That number is what justified building the citation gate, and it is what proved the gate worked when it went to 91%. No amount of faithfulness measurement would have surfaced it.
The work
Writing 180 pairs in two days
-
Mine real questions first, invent second
Search logs, issue trackers, the questions people actually ask in your team's chat. Sixty of mine came from real sources and they are noticeably harder than the ones I wrote from imagination, because real questions are phrased badly.
-
Write the question, then find the answer in the corpus
Never the reverse. Reading a passage and writing a question about it produces questions phrased in the passage's own words, which every retriever finds trivially. Your eval set then measures nothing and reports 0.97.
-
Record the source, not just the answer
The heading anchor or chunk id. This is a few extra seconds per pair and it is what makes the retrieval-versus-generation split possible later.
-
Deliberately vary the difficulty
Mine ended up roughly 40% single-passage, 30% spanning two sections, 20% unanswerable, 10% ambiguous-and-should-ask-back. An eval set of only easy questions has a ceiling you will hit in week one and learn nothing from after.
-
Version it and never quietly edit it
golden/v3.jsonl. When you change the set, bump the version — otherwise your metric history is comparing scores measured with different rulers, and the whole record becomes worthless.
Roughly seven minutes a pair. It is genuinely tedious and it is the highest-leverage two days I spent on the entire project.
Automation
Wiring it into CI, and picking the thresholds
An eval set you run when you remember to is a document. An eval set that blocks a merge is a gate. The difference is about fifteen lines of YAML.
on:
pull_request:
paths: ['src/**', 'prompts/**', 'golden/**'] # prompts count. They are code.
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pip install -r requirements.txt
- name: Run evaluation
run: python -m eval.run --dataset golden/v3.jsonl --concurrency 8
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
- name: Gate on regression
run: |
python -m eval.gate results.json \
--min-context-recall 0.80 \
--min-context-precision 0.75 \
--min-faithfulness 0.85 \
--min-refusal 0.85
- uses: actions/upload-artifact@v4
if: always() # you want the numbers MOST when it failed
with: { name: eval-results, path: results.json }
Two details in there are load-bearing. paths includes prompts/**, because a prompt edit changes behaviour as thoroughly as a code change and nothing else in CI will notice it. And if: always() on the artifact upload — the run you most want the detailed numbers from is the one that just failed, and a step that only uploads on success gives you a red X and nothing to read.
How to pick a threshold without making it up
I set mine wrong twice. Too high, and every PR fails and you start adding --skip-eval, at which point you have a gate that is off. Too low and it never fires.
-
Run the set ten times on unchanged code
Generation is stochastic, so the same system scores differently each run. Mine varied by ±0.018 on faithfulness. That spread is your noise floor.
-
Put the threshold below the noise, not below the mean
Mean 0.94, noise ±0.018 — a threshold at 0.93 fires on weather. I set 0.85: roughly four times the noise band below the current score, so it catches real movement and ignores dice.
-
Raise it when you improve, never lower it to pass
The threshold is a ratchet. The one time I lowered it to unblock myself, I spent the next week wondering whether the number meant anything, which it no longer did.
The receipts
The three pull requests it blocked
All three were mine. All three I believed in.
| The change | Why I made it | What actually happened | Outcome |
|---|---|---|---|
| Chunk size 800 → 400 tokens | Tighter passages, less noise in the prompt | Recall 0.89 → 0.71. Answers spanning two paragraphs were now split across chunks and only half was ever retrieved. | Reverted |
| Added "be concise" to the prompt | Answers were getting long | Faithfulness 0.94 → 0.82. Concision came out of the hedging and attribution, which is exactly the part citations depend on. | Reworded |
| Swapped to a faster reranker | 240 ms of p95 latency back | Recall 0.89 → 0.83. A genuine trade rather than a mistake — but a trade I would have made blind. | Kept the slow one |
The third one is the one I think about most. It was not a bug; it was a real latency-for-quality trade, and reasonable people could take either side. The gate did not make that decision for me. It just meant I was making it with two numbers instead of a feeling.
Honesty
What the eval set does not catch
This is where most write-ups stop, so let me be specific about the gaps, because a gate you over-trust is worse than one you know the shape of.
- Latency and cost. Nothing in those four metrics notices that a change doubled p95 or tripled the bill. Those are separate gates and I nearly forgot them.
- Tone. A change can make every answer correct, grounded, and unpleasant to read. No metric here objects.
- Adversarial input. My 180 questions are all asked in good faith. Prompt injection through a retrieved document is a completely different test suite, and I have not built it.
- Distribution drift. The set measures the questions I thought of in March. Real users ask new things; if the set never grows, it slowly stops describing production.
- Anything about the corpus being wrong. Faithfulness measures agreement with retrieved text. If the documentation is out of date, a perfectly faithful answer is a perfectly wrong one.
That last one is worth sitting with. Faithfulness is grounding, not truth. It answers "did the model make this up," never "is this correct" — and conflating the two is how a team ends up confidently serving stale documentation with a quality score attached.
The checklist
- A corpus you can personally grade — the prerequisite for everything else.
- 50 pairs minimum, 150+ before you trust a small movement.
- 20% unanswerable, phrased to sound exactly like the answerable ones.
- Expected sources on every pair, so retrieval and generation fail separately.
- Four metrics, not one score.
- Thresholds set from ten baseline runs, four noise-bands below the mean.
- The gate blocks the merge — a report nobody is forced to read is a report nobody reads.
- Prompts in the watched paths, versioned beside the code.
- Results uploaded on failure, not only on success.
- The set is versioned and grows from real questions that went wrong.
Where to start on Monday
Write twenty pairs. Not a hundred and eighty — twenty, this afternoon, including four that have no answer. Run them against what you have already built.
The number that comes back will be lower than you expect, and that gap between what you assumed and what you measured is the entire value of the exercise. Everything after it is just adding pairs.
- golden set
- context recall
- context precision
- faithfulness
- refusal correctness
- noise floor
- threshold ratchet
- prompts in CI
This is the measurement half of the four projects; the citation gate it justified is covered in the production RAG guide, and if you are gating an agent rather than a retrieval pipeline, the termination reasons in the agent loop post are the metrics that belong in a suite like this one. If you have found a metric that catches something these four miss, I would like to hear it — the comments are open.
Comments (0)
No comments yet
Be the first to share a thought on this article.
Join the conversation