Skip to content
A AhsanLab.Tech
AI 14 min read · March 27, 2026

Your AI System Is Failing Right Now and Nothing Is Alerting

On a Tuesday my RAG system got quietly worse. Every request returned 200, no alert fired, and the answers stayed fluent and confident while being subtly about the wrong thing. Why AI systems fail silently, what a trace must actually record, why the mean is a lie, and the forty-minute root cause that would have taken days.

A Ahsan Habib Save

On a Tuesday in week six my RAG system got worse. Not broken — worse. Every request returned 200. Nothing errored, nothing timed out, no alert fired, and the answers were still fluent, well-structured and confidently phrased. They were just subtly about the wrong thing, and I only noticed because I happened to ask it something I already knew the answer to.

Root cause took forty minutes. It would have taken days, because I know exactly what I would have done instead: I would have started rewriting prompts, since the prompt is the part you can see.

I know that because three weeks earlier I had done precisely that, for three days, on a problem that turned out to be retrieval.

3 dayslost, before tracingwrong layer entirely
40 minto root cause, aftersame class of bug
0alerts fired, both timesnothing was "down"

The shape of the problem

Silent failure is the default, not the exception

Normal backend monitoring is built on a premise that AI systems quietly violate: that failure produces a signal. A 500, a timeout, a queue backing up, a health check going red. Your dashboards are all downstream of something going wrong in a way a machine can notice.

An AI pipeline degrades without any of that happening:

What actually brokeWhat a normal monitor seesWhat the user gets
Retrieval returned off-topic chunks 200 OK, 1.4 s, normal A fluent answer about the wrong feature
An embedding model version changed 200 OK, no error rate movement Gradually worse answers, no cliff
A prompt edit dropped the citation instruction 200 OK, slightly fewer tokens Ungrounded claims, still well written
The model started hedging on everything 200 OK, slightly more tokens Useless non-answers, perfect grammar
An agent looped 40 times before answering 200 OK, 19 s — within timeout The right answer, at 20× the cost
Every row is a 200. That is the whole problem: your alerting is watching for a system that stops working, and this one keeps working while getting worse.

Instrumentation

Trace everything — and "everything" is specific

A trace here is not a log line. It is a structured record of every decision the pipeline made for one request, and it has to carry enough to reconstruct the request without re-running it. Concretely, per span:

ONE REQUEST, FULLY TRACED 1,412 ms TOTAL embed query 38 ms · 12 tok BM25 search 48 ms · 20 hits vector search 62 ms · 20 hits rerank 294 ms · 30 → 6 · top score 0.81 LLM generate 918 ms citation gate 52 ms chunk ids · rerank scores · the exact prompt · the response · 2,914 in / 186 out · $0.0031
FIGURE 1 — A SPAN WATERFALL · The rerank costs 21% of the request and the LLM 65%. You cannot know that from a single duration number.
rag/pipeline.py — instrumented
with trace.span("retrieve") as span:
    bm25 = keyword.search(query, k=20)
    dense = vectors.search(embedding, k=20)
    fused = rrf(bm25, dense)[:30]

    # Record the DECISION, not just the duration. When quality drops, the
    # question is always "what did it see", never "how long did it take".
    span.set(
        chunk_ids=[c.id for c in fused],
        bm25_top=[c.id for c in bm25[:3]],
        dense_top=[c.id for c in dense[:3]],
        overlap=len(set(bm25_ids) & set(dense_ids)),
    )

with trace.span("rerank") as span:
    top = reranker.rank(query, fused)[:6]
    span.set(
        kept=[c.id for c in top],
        scores=[round(c.score, 3) for c in top],
        dropped_best=round(fused[0].score, 3),   # did rerank overrule retrieval?
    )
Why self-hosted Langfuse Not ideology — usage ceilings. During the weeks when you are iterating hardest you generate the most traces, which is exactly when a managed free tier stops recording and you start sampling. Sampled traces are useless for "what happened to this request". It is one Docker Compose file and a Postgres volume, and I moved on with my life.

Aggregation

Four metrics, and never the mean

Traces answer "what happened to this request". Metrics answer "is it getting worse". You need both, and the metrics need choosing carefully:

MetricMineWhy this one
Latency p501.4 sWhat a typical user actually experiences
Latency p953.4 sWhat your angriest user experiences
Latency p996.1 sWhere the reranker queues behind itself
Cost per request$0.0031Multiply by projected traffic before promising anything
Citation coverage94.2%Proportion of answers fully grounded in retrieved text
Failure rate1.8%0.6% errors and 1.2% refusals — tracked separately

The mean is deliberately absent. The average of my p50 and p99 is a number describing nobody's experience, and it moves in the wrong direction: a change that makes 95% of requests faster and 5% catastrophically slower improves the mean.

p501.4 s
mean the lie1.9 s
p953.4 s
p996.1 s

Splitting errors from refusals matters just as much. An error is a bug. A refusal is the citation gate doing its job. Folded into one "failure rate", a system that starts refusing more — which might be correct behaviour after a corpus change — looks identical to one that started crashing.

The incident

The Tuesday, in full

This is the trace that paid for the whole layer.

  1. 11:20 — I notice, by accident

    I ask a question I know the answer to. The response is about a related feature. Plausible, well cited, wrong. Nothing on any dashboard is red.

  2. 11:26 — the metric confirms it is real

    Context recall over the last 24 hours: 0.61, down from 0.89. So it is not one unlucky question — it is systemic, and it started at some point I can now look for.

  3. 11:34 — the traces localise it

    I pull ten recent traces. Retrieval is returning chunks with reasonable similarity scores that are semantically unrelated to the query. So: not the prompt, not the model, not the reranker. The vectors themselves are wrong.

  4. 11:48 — the deploy marker names the change

    Recall steps down at a single deploy the previous evening. That deploy contained no retrieval code. It did contain a dependency update.

  5. 12:00 — root cause

    An unpinned sentence-transformers bumped the embedding model version. New queries were embedded by a different model than the index was built with, so every similarity score in the store was quietly meaningless. Nothing threw, because both models produce valid vectors of the same dimension.

CONTEXT RECALL · 14 DAYS NO ALERT FIRED 0.90 0.50 0.70 deploy · dependency bump no retrieval code changed pinned + reindexed four days at 0.61 · every request returned 200 OK
FIGURE 2 — THE DROP · Four days of degraded answers between the deploy and somebody noticing. The step is obvious in hindsight and invisible without the metric.
The detail that makes this generalise Both embedding models produced valid float vectors of identical dimension. There is no type error to catch, no exception to log, no schema to validate against. The only way this becomes visible is a quality metric measured over time. Every AI system has a class of bug shaped exactly like this one.

Closing the loop

From dashboard to gate

A dashboard tells you after. The point of building one is to work out what is worth measuring, and then to stop relying on somebody looking at it.

Dashboard only

Detection depends on a human opening a page and noticing a step change.

The Tuesday incident ran for four days before anyone did.

Dashboard + CI gate

The same metrics run against the golden set on every pull request, and a regression fails the build.

Detection moves from days to before merge.

The gate is covered properly in the eval set post; what matters here is the ordering. You cannot gate on a metric you have not defined, and you will not define the right metrics until traces have shown you which ones move when something breaks. Observability comes first and the gate falls out of it.

What to alert on, and what not to

I got this wrong initially by paging myself on quality metrics, which are noisy enough that I learned to ignore the notifications within a week — the worst possible outcome.

SignalTreatmentWhy
Error rate above baselineAlert immediatelyUnambiguous, actionable, rare
Cost per request doublingAlert immediatelyCheap to check, expensive to miss
p95 latency regressionAlert on sustained breachReal, but spiky enough to need a window
Citation coverage / recallGate in CI, review weeklyToo noisy per-request to page on; the gate catches the causes
Refusal rate climbingReview weeklyOften correct behaviour, occasionally a corpus problem

Honesty

What this layer costs

It is not free and the write-ups that pretend otherwise are selling something.

  • Storage grows fast. Full prompts and responses on every request is real volume. I sample non-error traces at 20% after fourteen days and keep every failure and refusal indefinitely.
  • Traces contain user text. The moment you record prompts you have built a datastore of whatever people typed. That needs the same retention and access thinking as any other user data, and it is easy to forget because it arrived as telemetry.
  • Instrumentation is code that rots. A span that stopped being set is invisible until the day you need it. Mine broke once when I refactored the retriever and nothing complained for two weeks.
  • Latency overhead is small but real — about 8 ms per request for me, async-flushed. Worth knowing rather than assuming zero.

The checklist

  • A span per pipeline stage, recording the decision and not only the duration.
  • Chunk ids and rerank scores in the trace — "what did it see" is the question you will always be asking.
  • The exact prompt sent, not a template reference.
  • Token counts in and out, and a dollar figure per request.
  • p50, p95 and p99 — and no mean anywhere.
  • Errors and refusals as separate metrics.
  • Deploy markers on every chart, or a step change names nothing.
  • Alert on errors and cost; gate on quality; review the rest weekly.
  • A retention policy, decided deliberately, because traces are user data.
  • Pin your model and embedding versions — the incident above is waiting in every unpinned dependency file.

Where to start on Monday

Instrument one pipeline, end to end, before you add another feature. Record the retrieved chunk ids if you record nothing else — that single field is what separates "the model is hallucinating" from "the model was handed the wrong paragraph", and those two have completely different fixes.

Then look at a hundred real traces. Not to find a bug — just to see what your system actually does, which in my experience is never quite what you believe it does.

  • Langfuse
  • span waterfall
  • p50 / p95 / p99
  • cost per request
  • citation coverage
  • deploy markers
  • trace retention
  • pinned versions

This is project three of the four and the one that looks weakest in a portfolio. It is also the one that found the bug. If you are instrumenting an agent rather than a pipeline, the termination reasons in the agent loop post are the first spans worth adding. Got a silent-failure story of your own? The comments are open — those are the ones worth collecting.

#Performance #AI Tools #LLM #Agents #CI/CD #RAG

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles

AI 13 min

Connecting OpenAI to a Real Backend: Lessons From SEO Automation

The naive version worked perfectly on fifty pages and failed four ways on twelve thousand — none of them about prompt quality. Wiring a model into a real backend is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle, and it needs everything batch processing always needed.

May 28, 2026
AI 15 min

The 940ms Voice Agent: Where Every Millisecond Goes

A conversation has a social latency budget, not a technical one — go past a second and people think the thing is broken. How a voice agent went from 1,530ms to 940ms to first audio, why every millisecond came from overlapping stages rather than speeding them up, the speculative optimisation that cost 38% more tokens for 60ms, and what a user hears when a dependency vanishes mid-sentence.

Apr 12, 2026