Your AI System Is Failing Right Now and Nothing Is Alerting
On a Tuesday my RAG system got quietly worse. Every request returned 200, no alert fired, and the answers stayed fluent and confident while being subtly about the wrong thing. Why AI systems fail silently, what a trace must actually record, why the mean is a lie, and the forty-minute root cause that would have taken days.
On a Tuesday in week six my RAG system got worse. Not broken — worse. Every request returned 200. Nothing errored, nothing timed out, no alert fired, and the answers were still fluent, well-structured and confidently phrased. They were just subtly about the wrong thing, and I only noticed because I happened to ask it something I already knew the answer to.
Root cause took forty minutes. It would have taken days, because I know exactly what I would have done instead: I would have started rewriting prompts, since the prompt is the part you can see.
I know that because three weeks earlier I had done precisely that, for three days, on a problem that turned out to be retrieval.
The shape of the problem
Silent failure is the default, not the exception
Normal backend monitoring is built on a premise that AI systems quietly violate: that failure produces a signal. A 500, a timeout, a queue backing up, a health check going red. Your dashboards are all downstream of something going wrong in a way a machine can notice.
An AI pipeline degrades without any of that happening:
| What actually broke | What a normal monitor sees | What the user gets |
|---|---|---|
| Retrieval returned off-topic chunks | 200 OK, 1.4 s, normal | A fluent answer about the wrong feature |
| An embedding model version changed | 200 OK, no error rate movement | Gradually worse answers, no cliff |
| A prompt edit dropped the citation instruction | 200 OK, slightly fewer tokens | Ungrounded claims, still well written |
| The model started hedging on everything | 200 OK, slightly more tokens | Useless non-answers, perfect grammar |
| An agent looped 40 times before answering | 200 OK, 19 s — within timeout | The right answer, at 20× the cost |
Instrumentation
Trace everything — and "everything" is specific
A trace here is not a log line. It is a structured record of every decision the pipeline made for one request, and it has to carry enough to reconstruct the request without re-running it. Concretely, per span:
with trace.span("retrieve") as span:
bm25 = keyword.search(query, k=20)
dense = vectors.search(embedding, k=20)
fused = rrf(bm25, dense)[:30]
# Record the DECISION, not just the duration. When quality drops, the
# question is always "what did it see", never "how long did it take".
span.set(
chunk_ids=[c.id for c in fused],
bm25_top=[c.id for c in bm25[:3]],
dense_top=[c.id for c in dense[:3]],
overlap=len(set(bm25_ids) & set(dense_ids)),
)
with trace.span("rerank") as span:
top = reranker.rank(query, fused)[:6]
span.set(
kept=[c.id for c in top],
scores=[round(c.score, 3) for c in top],
dropped_best=round(fused[0].score, 3), # did rerank overrule retrieval?
)
Aggregation
Four metrics, and never the mean
Traces answer "what happened to this request". Metrics answer "is it getting worse". You need both, and the metrics need choosing carefully:
| Metric | Mine | Why this one |
|---|---|---|
| Latency p50 | 1.4 s | What a typical user actually experiences |
| Latency p95 | 3.4 s | What your angriest user experiences |
| Latency p99 | 6.1 s | Where the reranker queues behind itself |
| Cost per request | $0.0031 | Multiply by projected traffic before promising anything |
| Citation coverage | 94.2% | Proportion of answers fully grounded in retrieved text |
| Failure rate | 1.8% | 0.6% errors and 1.2% refusals — tracked separately |
The mean is deliberately absent. The average of my p50 and p99 is a number describing nobody's experience, and it moves in the wrong direction: a change that makes 95% of requests faster and 5% catastrophically slower improves the mean.
Splitting errors from refusals matters just as much. An error is a bug. A refusal is the citation gate doing its job. Folded into one "failure rate", a system that starts refusing more — which might be correct behaviour after a corpus change — looks identical to one that started crashing.
The incident
The Tuesday, in full
This is the trace that paid for the whole layer.
-
11:20 — I notice, by accident
I ask a question I know the answer to. The response is about a related feature. Plausible, well cited, wrong. Nothing on any dashboard is red.
-
11:26 — the metric confirms it is real
Context recall over the last 24 hours: 0.61, down from 0.89. So it is not one unlucky question — it is systemic, and it started at some point I can now look for.
-
11:34 — the traces localise it
I pull ten recent traces. Retrieval is returning chunks with reasonable similarity scores that are semantically unrelated to the query. So: not the prompt, not the model, not the reranker. The vectors themselves are wrong.
-
11:48 — the deploy marker names the change
Recall steps down at a single deploy the previous evening. That deploy contained no retrieval code. It did contain a dependency update.
-
12:00 — root cause
An unpinned
sentence-transformersbumped the embedding model version. New queries were embedded by a different model than the index was built with, so every similarity score in the store was quietly meaningless. Nothing threw, because both models produce valid vectors of the same dimension.
Closing the loop
From dashboard to gate
A dashboard tells you after. The point of building one is to work out what is worth measuring, and then to stop relying on somebody looking at it.
Detection depends on a human opening a page and noticing a step change.
The Tuesday incident ran for four days before anyone did.
The same metrics run against the golden set on every pull request, and a regression fails the build.
Detection moves from days to before merge.
The gate is covered properly in the eval set post; what matters here is the ordering. You cannot gate on a metric you have not defined, and you will not define the right metrics until traces have shown you which ones move when something breaks. Observability comes first and the gate falls out of it.
What to alert on, and what not to
I got this wrong initially by paging myself on quality metrics, which are noisy enough that I learned to ignore the notifications within a week — the worst possible outcome.
| Signal | Treatment | Why |
|---|---|---|
| Error rate above baseline | Alert immediately | Unambiguous, actionable, rare |
| Cost per request doubling | Alert immediately | Cheap to check, expensive to miss |
| p95 latency regression | Alert on sustained breach | Real, but spiky enough to need a window |
| Citation coverage / recall | Gate in CI, review weekly | Too noisy per-request to page on; the gate catches the causes |
| Refusal rate climbing | Review weekly | Often correct behaviour, occasionally a corpus problem |
Honesty
What this layer costs
It is not free and the write-ups that pretend otherwise are selling something.
- Storage grows fast. Full prompts and responses on every request is real volume. I sample non-error traces at 20% after fourteen days and keep every failure and refusal indefinitely.
- Traces contain user text. The moment you record prompts you have built a datastore of whatever people typed. That needs the same retention and access thinking as any other user data, and it is easy to forget because it arrived as telemetry.
- Instrumentation is code that rots. A span that stopped being set is invisible until the day you need it. Mine broke once when I refactored the retriever and nothing complained for two weeks.
- Latency overhead is small but real — about 8 ms per request for me, async-flushed. Worth knowing rather than assuming zero.
The checklist
- A span per pipeline stage, recording the decision and not only the duration.
- Chunk ids and rerank scores in the trace — "what did it see" is the question you will always be asking.
- The exact prompt sent, not a template reference.
- Token counts in and out, and a dollar figure per request.
- p50, p95 and p99 — and no mean anywhere.
- Errors and refusals as separate metrics.
- Deploy markers on every chart, or a step change names nothing.
- Alert on errors and cost; gate on quality; review the rest weekly.
- A retention policy, decided deliberately, because traces are user data.
- Pin your model and embedding versions — the incident above is waiting in every unpinned dependency file.
Where to start on Monday
Instrument one pipeline, end to end, before you add another feature. Record the retrieved chunk ids if you record nothing else — that single field is what separates "the model is hallucinating" from "the model was handed the wrong paragraph", and those two have completely different fixes.
Then look at a hundred real traces. Not to find a bug — just to see what your system actually does, which in my experience is never quite what you believe it does.
- Langfuse
- span waterfall
- p50 / p95 / p99
- cost per request
- citation coverage
- deploy markers
- trace retention
- pinned versions
This is project three of the four and the one that looks weakest in a portfolio. It is also the one that found the bug. If you are instrumenting an agent rather than a pipeline, the termination reasons in the agent loop post are the first spans worth adding. Got a silent-failure story of your own? The comments are open — those are the ones worth collecting.
Comments (0)
No comments yet
Be the first to share a thought on this article.
Join the conversation