Skip to content
A AhsanLab.Tech
AI 14 min read · February 24, 2026

How Do You Know Your AI Got Worse? Build the Eval Set First

I opened a pull request that made my RAG system measurably worse and was completely certain it was an improvement. The build failed instead. What a golden evaluation pair actually contains, the four metrics that localise a failure, why 21% of my questions have no answer, how to pick a threshold from the noise floor, and the three of my own PRs it blocked.

A Ahsan Habib Save

I opened a pull request that made my RAG system meaningfully worse, and I was completely certain it was an improvement. Smaller chunks, tighter passages, less noise in the prompt — every instinct I had said yes. I spot-checked six questions, they looked fine, and I would have merged it on a Tuesday afternoon without a second thought.

The build failed instead. Context recall had gone from 0.89 to 0.71, because a chunk small enough to be "precise" was too small to contain a whole answer, and half my corpus answers span two paragraphs. Six spot-checks did not find that. A hundred and eighty did.

This post is about the least glamorous artefact in an AI system and the one I would rebuild first if I lost everything: the evaluation set, and the CI gate that makes it mean something.

180hand-verified pairstwo days of work
38deliberately unanswerablethe 21% that matter most
3of my own PRs blockedin six weeks

Why intuition fails here specifically

Every AI change feels like an improvement

In ordinary backend work, a regression announces itself. The test goes red, the type checker complains, the endpoint 500s. You do not need discipline to notice — you need discipline to fix it.

AI systems have the opposite shape. A regression here produces a fluent, well-structured, confident answer that happens to be worse. Nothing throws. The output still looks exactly like output. And because generation is stochastic, running the same question twice gives you two different answers, so your instinct to "just check it again" produces noise rather than signal.

A regression in a normal system breaks something. A regression in an AI system produces a slightly worse paragraph, and paragraphs all look the same.

So spot-checking does not scale, and it does not even work at small scale, for a reason worth stating plainly: you pick the questions you already thought about. Those are the questions your last change was designed to fix.

The artefact

What a golden pair actually contains

Not a question and an answer. Four things:

golden/v3.jsonl — one line per pair
{
  "id": "elo-014",
  "question": "How do I stop a global scope applying to one query?",
  "ground_truth": "Call withoutGlobalScope(ScopeClass::class) on the query,
                   or withoutGlobalScopes() to remove all of them.",
  "expected_sources": ["eloquent/global-scopes#removing"],
  "answerable": true,
  "note": "phrased as intent, not as the API name — tests semantic retrieval"
}
{
  "id": "neg-007",
  "question": "What is the default connection pool size in Filament?",
  "ground_truth": null,
  "expected_sources": [],
  "answerable": false,
  "note": "sounds plausible, has no answer. Filament has no pool."
}

expected_sources is what separates an eval set from a quiz. It lets you ask did retrieval find the right passage separately from did the model use it well — and those two failures need completely different fixes. Without it, every regression looks the same and you tune the prompt because the prompt is the thing you can see.

The prerequisite nobody mentions You must be able to grade the corpus yourself. My first attempt used research papers in a field I do not work in; I could not tell a correct answer from a confident one, so every "ground truth" I wrote was a guess and the whole set measured nothing. I started over on documentation I use daily. Pick a corpus where you are the expert, or you are building a ruler out of guesses.
GOLDEN SET 180 pairs · versioned PULL REQUEST code OR prompt RUN EVAL 2.5 min · $0.42 FOUR METRICS not one score GATE thresholds MERGEABLE all four pass BUILD FAILS any one below the change does not merge — fix it and push again the same script runs locally
FIGURE 1 — THE GATE · A prompt edit goes through this too. It is a behaviour change with no diff in the type checker.

Measurement

Four metrics, because one number cannot answer four questions

A single "quality score" is a comfortable lie. When it drops you know something is wrong and nothing about where. These four localise the failure:

MetricThe question it answersFails whenMy gate
Context recall Did retrieval find the right passage at all? Chunking, embeddings, or the query itself ≥ 0.80
Context precision Is the right passage near the top? The reranker, or the fusion step ≥ 0.75
Faithfulness Are the claims supported by what was retrieved? The prompt, or the model filling gaps ≥ 0.85
Refusal correctness Does it decline when there is no answer? The citation gate, or its absence ≥ 0.85

The pairing matters more than the individual numbers. High faithfulness with low recall means the system is honestly summarising the wrong passage. High recall with low faithfulness means it found the answer and then embroidered it. Those look identical on a single score and need opposite fixes.

Recall chunk PR0.71
Recall reverted0.89
Faithfulness "be concise"0.82
Faithfulness reverted0.94

Two blocked pull requests, two different metrics, two entirely unrelated causes. A single score would have told me "it got worse" twice and pointed at nothing.

The part that changed the system

The 21% of questions with no answer

Thirty-eight of my 180 pairs are unanswerable on purpose. They sound completely reasonable, they use the right vocabulary, and the corpus simply does not cover them. Writing them felt like padding. They turned out to be the most valuable part of the set by a wide margin.

Without them

Every metric looked healthy. Faithfulness 0.88, recall 0.89.

The system invented a confident answer for 76% of questions it had no business answering — and nothing in my dashboard could see it, because a hallucinated answer to a question with no ground truth is not a wrong answer. It is an unmeasured one.

With them

Refusal correctness appeared as a number: 24%.

That number is what justified building the citation gate, and it is what proved the gate worked when it went to 91%. No amount of faithfulness measurement would have surfaced it.

The trap in generating these Do not ask a model to write your unanswerable questions. I tried; it produced questions that were obviously unanswerable — different vocabulary, off-topic, easy. The useful negatives are the ones that sound exactly like the positives, and you only find those by knowing the corpus well enough to know its edges.

The work

Writing 180 pairs in two days

  1. Mine real questions first, invent second

    Search logs, issue trackers, the questions people actually ask in your team's chat. Sixty of mine came from real sources and they are noticeably harder than the ones I wrote from imagination, because real questions are phrased badly.

  2. Write the question, then find the answer in the corpus

    Never the reverse. Reading a passage and writing a question about it produces questions phrased in the passage's own words, which every retriever finds trivially. Your eval set then measures nothing and reports 0.97.

  3. Record the source, not just the answer

    The heading anchor or chunk id. This is a few extra seconds per pair and it is what makes the retrieval-versus-generation split possible later.

  4. Deliberately vary the difficulty

    Mine ended up roughly 40% single-passage, 30% spanning two sections, 20% unanswerable, 10% ambiguous-and-should-ask-back. An eval set of only easy questions has a ceiling you will hit in week one and learn nothing from after.

  5. Version it and never quietly edit it

    golden/v3.jsonl. When you change the set, bump the version — otherwise your metric history is comparing scores measured with different rulers, and the whole record becomes worthless.

Roughly seven minutes a pair. It is genuinely tedious and it is the highest-leverage two days I spent on the entire project.

Automation

Wiring it into CI, and picking the thresholds

An eval set you run when you remember to is a document. An eval set that blocks a merge is a gate. The difference is about fifteen lines of YAML.

.github/workflows/eval.yml
on:
  pull_request:
    paths: ['src/**', 'prompts/**', 'golden/**']   # prompts count. They are code.

jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pip install -r requirements.txt

      - name: Run evaluation
        run: python -m eval.run --dataset golden/v3.jsonl --concurrency 8
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

      - name: Gate on regression
        run: |
          python -m eval.gate results.json \
            --min-context-recall    0.80 \
            --min-context-precision 0.75 \
            --min-faithfulness      0.85 \
            --min-refusal           0.85

      - uses: actions/upload-artifact@v4
        if: always()                    # you want the numbers MOST when it failed
        with: { name: eval-results, path: results.json }

Two details in there are load-bearing. paths includes prompts/**, because a prompt edit changes behaviour as thoroughly as a code change and nothing else in CI will notice it. And if: always() on the artifact upload — the run you most want the detailed numbers from is the one that just failed, and a step that only uploads on success gives you a red X and nothing to read.

How to pick a threshold without making it up

I set mine wrong twice. Too high, and every PR fails and you start adding --skip-eval, at which point you have a gate that is off. Too low and it never fires.

  1. Run the set ten times on unchanged code

    Generation is stochastic, so the same system scores differently each run. Mine varied by ±0.018 on faithfulness. That spread is your noise floor.

  2. Put the threshold below the noise, not below the mean

    Mean 0.94, noise ±0.018 — a threshold at 0.93 fires on weather. I set 0.85: roughly four times the noise band below the current score, so it catches real movement and ignores dice.

  3. Raise it when you improve, never lower it to pass

    The threshold is a ratchet. The one time I lowered it to unblock myself, I spent the next week wondering whether the number meant anything, which it no longer did.

FAITHFULNESS PER PULL REQUEST SIX WEEKS · 28 PRs 0.85 1.00 0.60 chunks 800→400 · 0.71 "be concise" · 0.82 faster reranker · 0.83
FIGURE 2 — SIX WEEKS OF THE GATE · Three red bars. Each one felt like an improvement while I was writing it.

The receipts

The three pull requests it blocked

All three were mine. All three I believed in.

The changeWhy I made itWhat actually happenedOutcome
Chunk size 800 → 400 tokens Tighter passages, less noise in the prompt Recall 0.89 → 0.71. Answers spanning two paragraphs were now split across chunks and only half was ever retrieved. Reverted
Added "be concise" to the prompt Answers were getting long Faithfulness 0.94 → 0.82. Concision came out of the hedging and attribution, which is exactly the part citations depend on. Reworded
Swapped to a faster reranker 240 ms of p95 latency back Recall 0.89 → 0.83. A genuine trade rather than a mistake — but a trade I would have made blind. Kept the slow one

The third one is the one I think about most. It was not a bug; it was a real latency-for-quality trade, and reasonable people could take either side. The gate did not make that decision for me. It just meant I was making it with two numbers instead of a feeling.

What this buys beyond catching regressions Once the gate exists, you can be aggressive. I tried nine retrieval changes in the following fortnight and merged six, quickly, without agonising over any of them — because the cost of being wrong dropped to a failed build. A test suite makes you brave for the same reason.

Honesty

What the eval set does not catch

This is where most write-ups stop, so let me be specific about the gaps, because a gate you over-trust is worse than one you know the shape of.

  • Latency and cost. Nothing in those four metrics notices that a change doubled p95 or tripled the bill. Those are separate gates and I nearly forgot them.
  • Tone. A change can make every answer correct, grounded, and unpleasant to read. No metric here objects.
  • Adversarial input. My 180 questions are all asked in good faith. Prompt injection through a retrieved document is a completely different test suite, and I have not built it.
  • Distribution drift. The set measures the questions I thought of in March. Real users ask new things; if the set never grows, it slowly stops describing production.
  • Anything about the corpus being wrong. Faithfulness measures agreement with retrieved text. If the documentation is out of date, a perfectly faithful answer is a perfectly wrong one.

That last one is worth sitting with. Faithfulness is grounding, not truth. It answers "did the model make this up," never "is this correct" — and conflating the two is how a team ends up confidently serving stale documentation with a quality score attached.

The checklist

  • A corpus you can personally grade — the prerequisite for everything else.
  • 50 pairs minimum, 150+ before you trust a small movement.
  • 20% unanswerable, phrased to sound exactly like the answerable ones.
  • Expected sources on every pair, so retrieval and generation fail separately.
  • Four metrics, not one score.
  • Thresholds set from ten baseline runs, four noise-bands below the mean.
  • The gate blocks the merge — a report nobody is forced to read is a report nobody reads.
  • Prompts in the watched paths, versioned beside the code.
  • Results uploaded on failure, not only on success.
  • The set is versioned and grows from real questions that went wrong.

Where to start on Monday

Write twenty pairs. Not a hundred and eighty — twenty, this afternoon, including four that have no answer. Run them against what you have already built.

The number that comes back will be lower than you expect, and that gap between what you assumed and what you measured is the entire value of the exercise. Everything after it is just adding pairs.

  • golden set
  • context recall
  • context precision
  • faithfulness
  • refusal correctness
  • noise floor
  • threshold ratchet
  • prompts in CI

This is the measurement half of the four projects; the citation gate it justified is covered in the production RAG guide, and if you are gating an agent rather than a retrieval pipeline, the termination reasons in the agent loop post are the metrics that belong in a suite like this one. If you have found a metric that catches something these four miss, I would like to hear it — the comments are open.

#AI Tools #LLM #Agents #CI/CD #RAG #Prompt Engineering

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles

AI 13 min

Connecting OpenAI to a Real Backend: Lessons From SEO Automation

The naive version worked perfectly on fifty pages and failed four ways on twelve thousand — none of them about prompt quality. Wiring a model into a real backend is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle, and it needs everything batch processing always needed.

May 28, 2026
AI 15 min

The 940ms Voice Agent: Where Every Millisecond Goes

A conversation has a social latency budget, not a technical one — go past a second and people think the thing is broken. How a voice agent went from 1,530ms to 940ms to first audio, why every millisecond came from overlapping stages rather than speeding them up, the speculative optimisation that cost 38% more tokens for 60ms, and what a user hears when a dependency vanishes mid-sentence.

Apr 12, 2026
AI 14 min

Your AI System Is Failing Right Now and Nothing Is Alerting

On a Tuesday my RAG system got quietly worse. Every request returned 200, no alert fired, and the answers stayed fluent and confident while being subtly about the wrong thing. Why AI systems fail silently, what a trace must actually record, why the mean is a lie, and the forty-minute root cause that would have taken days.

Mar 27, 2026