Skip to content
A AhsanLab.Tech
AI 18 min read · January 25, 2026

I Built Four AI Engineering Projects in 12 Weeks — Here Is Every Number

A production RAG system with enforced citations, a local model benchmark, an observability layer, and a real-time voice agent — built in sequence as one system, measured end to end. Why I deliberately skipped fine-tuning, the numbers from every phase, and the four things I got wrong first.

A Ahsan Habib Save

Twelve weeks ago I had the same problem every engineer moving into AI has: a folder full of tutorial chatbots that all worked perfectly on the three questions I had thought to ask them, and no idea whether any of it resembled the work an actual AI team does.

So I picked four projects, gave myself twelve weeks, and made one rule: every claim has to be a number I measured myself. No "significantly faster." No "much more accurate." A number, from a script anyone can re-run.

This is the write-up. Four projects, what each one taught me, the numbers that came out of them, and the three days I lost to a bug I would never have found without the boring project nobody builds.

First: what I deliberately did not build

Most "AI portfolio" lists put fine-tuning in the middle — LoRA, QLoRA, DPO, training curves. I dropped it on purpose, and I want to explain why before anything else, because the reasoning shaped every other decision here.

There is a line between two jobs that get advertised under one name:

An AI engineer composes models into reliable systems. An ML engineer changes the model's weights.

Retrieval, tool calling, constrained generation, evaluation, observability, streaming, failure handling — that is composition, and that is the job I want. Fine-tuning is a genuinely valuable skill, but it is a different one, and a portfolio that gestures weakly at both says less than one that is unmistakably deep in one.

There is also an honest engineering argument. Fine-tuning is what you reach for after you have demonstrated a ceiling — after you can show, with a golden evaluation set and real metrics, that careful prompting and better retrieval cannot get you there. I did not have that evidence twelve weeks ago. Nobody starting out does. Building the measurement apparatus first is not a detour around fine-tuning; it is the thing that tells you whether you need it.

If you are choosing between these tracks: build the composition layer first regardless. An ML engineer who cannot evaluate a system is stuck, and an AI engineer who can evaluate one is immediately useful. The eval set is the prerequisite for both.

The four projects are one system

I did not build four unrelated repos. Each project is a layer of the same thing, and the order matters: project 1 produces the evaluation set that project 3 gates on, and project 2 produces the local model that project 4 falls back to when the cloud times out.

REQUEST PATH — PROJECTS 1, 2 & 4 VOICE IN ASR · streaming AGENT LOOP decide · call tools RETRIEVE BM25 + vector RERANK cross-encoder LLM + citation gate VOICE OUT TTS · streaming LOCAL SLM Llama 3.2 3B timeout → PROJECT 3 — TRACING & METRICS (Langfuse, self-hosted) every chunk retrieved · every rerank score · p50 / p95 · cost per request · citation coverage CI REGRESSION GATE — 180 golden Q/A pairs on every pull request faithfulness < 0.85 or context recall < 0.80 → the build fails and the change does not merge
FIGURE 1 — Four projects, one system. The observability band and the CI gate sit under everything, which is exactly why they were the hardest part to skip.

Project 1 — A RAG system that refuses to answer

The corpus is Laravel 12, Filament and Livewire documentation: 1,240 pages, 8,900 chunks. I picked it because I work in it daily, which means I can judge whether an answer is right. That sounds obvious and it is the single most important decision in the project — a golden evaluation set built on a corpus you cannot personally grade is a set of guesses in a CSV.

Phase 1 — the forty-line version

Chunk at 500–800 tokens with roughly 100 tokens of overlap, embed, store in Chroma, retrieve top-5, stuff into the prompt, ask for citations. It works, it is genuinely satisfying, and it is where most portfolio projects stop.

Here is what it scored on my evaluation set:

Configuration Faithfulness Context recall Refusal correctness p95 latency
Dense vector only, top-5 0.71 0.68 12% 1.9 s
+ BM25 hybrid, RRF fusion 0.79 0.84 18% 2.1 s
+ cross-encoder rerank 0.88 0.89 24% 3.4 s
+ citation enforcement gate 0.94 0.89 91% 3.5 s

Read the refusal correctness column, not the faithfulness one. That is the percentage of the 38 deliberately unanswerable questions in my set where the system correctly said "the documentation does not cover this" instead of inventing something. The naive pipeline got that right 12% of the time. It confidently made up an answer to almost nine out of ten questions it had no business answering — and every one of those answers was well-formatted, plausible, and wrong.

That number is the whole reason this project exists.

Phase 2 — hybrid retrieval and the rerank

QUESTION from the user BM25 KEYWORD exact terms, top-20 DENSE VECTOR meaning, top-20 RRF FUSION → 30 candidates CROSS-ENCODER rescore → top-6 CITATION GATE supported? ANSWER every claim cited REFUSAL "not covered" yes no
FIGURE 2 — HYBRID RETRIEVAL · Two retrievers disagree usefully. BM25 catches withoutGlobalScopes; the vector search catches "how do I stop a scope applying." Reciprocal Rank Fusion needs no score normalisation, which is why it beat the weighted-sum blend I tried first.

The fusion step is the part people over-engineer. I started with a weighted sum of normalised scores and spent a day tuning the weight, then threw it away for Reciprocal Rank Fusion, which ignores scores entirely and only looks at rank position:

# retrieval/fuse.py — RRF needs no score normalisation, which is the point.
# BM25 scores and cosine similarities live on incomparable scales; ranks don't.

def reciprocal_rank_fusion(rankings: list[list[str]], k: int = 60) -> list[str]:
    scores: dict[str, float] = {}
    for ranking in rankings:
        for position, chunk_id in enumerate(ranking):
            scores[chunk_id] = scores.get(chunk_id, 0.0) + 1.0 / (k + position + 1)
    return sorted(scores, key=scores.get, reverse=True)

candidates = reciprocal_rank_fusion([
    bm25.search(query, top_k=20),
    vectors.search(embed(query), top_k=20),
])[:30]

# The reranker is the expensive step, so it runs once, on 30 candidates, not 8,900.
top_chunks = reranker.rank(query, candidates)[:6]

The reranker costs 380 ms at p95 and moved faithfulness from 0.79 to 0.88. That is the best latency-for-quality trade in the entire system, and it is one model call.

The citation gate

This is the piece I would keep if I could keep only one. Before an answer is returned, every sentence in it is checked against the retrieved chunks. If a claim is not supported, the answer does not go out.

# answer/gate.py — a refusal is a feature, not an error path.

result = llm.generate(prompt, chunks=top_chunks)

unsupported = [
    claim for claim in split_claims(result.text)
    if not any(supports(chunk, claim) for chunk in top_chunks)
]

if unsupported:
    log.info("gate.refused", claims=len(unsupported), query=query)
    return Refusal(
        message="The documentation I have does not cover this.",
        nearest=top_chunks[:2],   # show what we DID find — never a dead end
    )

return Answer(text=result.text, citations=result.citations)

Refusal correctness went from 24% to 91%. The remaining 9% are questions where the docs contain something adjacent enough that the check passes — a partial answer to a question about a related feature. I have not solved those and I am not going to pretend otherwise.

Phase 3 — the golden set and the CI gate

180 question/answer pairs, hand-verified, 38 of them deliberately unanswerable. It took two full days to build and it is the most valuable artefact in all four projects. Ragas runs it on every pull request; if faithfulness drops below 0.85 or context recall below 0.80, the build fails.

In six weeks that gate blocked three of my own pull requests. One was a chunking change I was certain was an improvement. It was not.

Project 2 — Running it all offline on a laptop

The question I kept getting asked: why bother with a 3B model when a frontier API is one HTTP call away? The answer is that "one HTTP call away" assumes a network, a budget, and permission to send the data — and in real work at least one of those three is regularly missing. Privacy rules, latency floors, cost at scale, and edge deployments with no connectivity are not edge cases; they are Tuesday.

Ollama, a FastAPI wrapper, and a benchmark harness that writes CSV rather than screenshots. Everything below is on one machine — Ryzen 7 5800H, 32 GB RAM, RTX 3060 6 GB — over the same 30-prompt suite, five runs each.

Model Quant Memory Time to first token Tokens/sec Valid JSON
Llama 3.2 3B Q4_K_M 2.4 GB 180 ms 61.3 96.7%
Phi-4-mini 3.8B Q4_K_M 2.9 GB 210 ms 52.8 98.3%
Mistral 7B Q4_K_M 4.6 GB 340 ms 28.4 91.2%
Mistral 7B Q5_K_M 5.4 GB 390 ms 22.1 93.8%

The Q4 → Q5 row is the one worth staring at. Going from Q4_K_M to Q5_K_M cost 22% of throughput and 800 MB of memory and bought 2.6 percentage points of schema compliance. On this hardware, for this task, that is a bad trade — and I would not have known that from any blog post, because it depends entirely on the task and the machine.

Phi-4-mini won on the metric I actually cared about, despite losing on the one everyone quotes. That is the entire lesson of the project.

Structure, validation, retry

Raw schema violation rate was 8.4%. One retry with the validation error fed back into the prompt took it to 0.9%:

# local/structured.py — constrain, validate, retry once, then fail honestly.

class Extraction(BaseModel):
    intent: Literal["question", "command", "chitchat"]
    entities: list[str]
    confidence: float = Field(ge=0.0, le=1.0)

def extract(text: str, retries: int = 1) -> Extraction | None:
    prompt = EXTRACT_PROMPT.format(text=text)
    for attempt in range(retries + 1):
        raw = ollama.generate(prompt, format="json", temperature=0.0)
        try:
            return Extraction.model_validate_json(raw)
        except ValidationError as e:
            if attempt == retries:
                log.warning("extract.failed", text=text[:80], error=str(e))
                return None                      # fail visibly, never silently
            prompt = REPAIR_PROMPT.format(raw=raw, error=e)

Note retries: int = 1. Not 3, not a while True. A model that produced garbage twice with the error handed back to it is not going to produce gold on the fifth attempt — it is going to produce a latency spike and a bill.

The temperature study was the cheapest experiment here and the one that changed how I think. Same 30 prompts, five runs each: at temperature=0, 3 of 30 prompts produced any variation at all across runs. At 0.7, 27 of 30 did. If you are extracting structured data and your temperature is not zero, you have chosen variance for no reason.

Project 3 — The boring one that found the bug

This is the project that looks worst in a portfolio and taught me the most. No pretty frontend. No demo you can show someone in ten seconds. Just tracing on every step of the RAG pipeline, four metrics tracked over time, and a CI gate.

I self-hosted Langfuse rather than using a managed tier, purely so I would not hit a usage ceiling while iterating. Every request records which chunks were retrieved, how the reranker reordered them, the exact prompt sent, the response, and the token counts.

Metric Value Why this one
Latency p50 1.4 s What a typical user feels
Latency p95 3.4 s What your angriest user feels
Latency p99 6.1 s Where the reranker queues up
Cost per request $0.0031 Multiply by traffic before promising anything
Citation coverage 94.2% Proportion of answers fully grounded
Failure rate 1.8% 0.6% errors, 1.2% refusals — tracked separately

Note the mean is nowhere on that table. The average of 1.4 s and 6.1 s is a number that describes nobody's experience.

The Tuesday that justified the whole project

Six weeks in, answer quality got noticeably worse. Not broken — worse. Nothing errored, nothing alerted, the system returned confident answers all day. If I had been running the naive setup I would have assumed I was imagining it.

Instead I opened the dashboard, found context recall had dropped from 0.89 to 0.61, and used the traces to see that the retrieved chunks had gone subtly off-topic. Root cause: an unpinned dependency had bumped the embedding model version. The new vectors and the ones already in the index were computed by different models, so every similarity score in the store was quietly meaningless.

Forty minutes from noticing to root cause, because the traces existed. Without them I would have spent days rewriting prompts that were never the problem — and I know that, because that is exactly what I did for three days earlier in the project, before this layer existed.

# .github/workflows/eval.yml — a prompt change is a behaviour change.
- name: RAG evaluation
  run: python -m eval.run --dataset golden/v3.jsonl --out results.json

- name: Gate on regression
  run: |
    python -m eval.gate results.json \
      --min-faithfulness 0.85 \
      --min-context-recall 0.80
    # exits non-zero → the build fails → the change does not merge

Prompts live in prompts/*.yaml, versioned next to the code and diffed in review. A one-word prompt edit can move faithfulness three points. Treating that as a config change rather than a code change is how teams ship regressions they cannot explain.

Project 4 — Voice, where every millisecond is visible

I picked the voice track over the vision and log-analysis options because a conversation is unforgiving about latency in a way a dashboard never is. Nobody notices 400 ms in an API. Everybody notices it in a reply.

Mic → ASR → retrieval → LLM → TTS → speaker, orchestrated over WebSockets with a structured event protocol. Phase one was just getting audio to flow end to end, which took longer than I expected and optimised nothing.

Phase two was the interesting part: decomposing the latency instead of measuring it as one number.

TIME TO FIRST AUDIO — p50 BREAKDOWN, BEFORE OPTIMISATION 0 ms 500 ms 1000 ms 1500 ms VAD end-of-speech 240 ms ASR finalisation 310 ms Retrieval + rerank 290 ms LLM time to first token 420 ms — the biggest slice TTS time to first byte 180 ms Orchestration + network 940 ms — after optimisation
FIGURE 3 — LATENCY BUDGET · You cannot optimise "1.5 seconds." You can optimise a 420 ms LLM time-to-first-token, and you can only see it once the total is decomposed per request.

Total time to first audio started at 1,530 ms and finished at 940 ms (p95: 1.8 s). Three changes did nearly all of it:

  • Start retrieving on partial transcripts rather than waiting for ASR to finalise. Overlapping those two stages saved ~260 ms and cost some wasted retrievals, which is a trade I will take every time.
  • Chunk TTS by sentence. Speaking the first sentence while the model is still generating the second removes the entire tail of generation from perceived latency.
  • Keep the reranker warm. A cold cross-encoder added 600 ms to the first request after any idle period, which is exactly the request a person is most likely to be judging.

Phase 3 — what happens when things break

This is where project 2 stopped being a separate repo and became a component. Every dependency gets a timeout and a documented fallback:

Failure Detected by Degradation Cost
ASR unavailable WebSocket close / 3 s timeout Switch to typed input, say so out loud Voice input lost, session survives
Cloud LLM timeout 4 s deadline Fall back to local Llama 3.2 3B +260 ms, shorter answers
TTS unavailable 5 s timeout Return the answer as text No audio, answer still delivered
Retrieval empty Zero chunks over threshold Refuse, offer to rephrase No answer — correctly

The rule underneath all four rows: the system never hangs silently. It degrades visibly and says what it is doing. A voice assistant that goes quiet for eight seconds has failed more completely than one that says "let me try that a simpler way."

Replay mode was the last thing I built and the thing I would build first next time. 74 recorded sessions can be fed back through the pipeline deterministically, which turns "it did something weird once yesterday" from a shrug into a test case.

What I would tell someone starting this on Monday

Build the evaluation set before the second feature. Not after — before. Every improvement I made after week five was measurable, and every improvement I made before it was a guess I felt good about. The two days spent hand-writing 180 question/answer pairs bought more than any other two days in the twelve weeks.

Four other things I got wrong first:

  • I optimised prompts for three days when the problem was retrieval. The model was doing fine with what it was handed; what it was handed was wrong. Without traces I could not see that, so I tuned the one thing I could see.
  • I picked a corpus I could not grade. My first attempt used a research-paper set in a field I do not know. I could not tell a good answer from a confident one, which made the whole evaluation set worthless. Starting over on documentation I use daily was the right call.
  • I reached for a framework too early. The first agent loop was LangGraph wrapping a single model call with two tools. The model SDK's own tool-runner did the same job in a third of the code. Frameworks earn their place when you have several coordinating agents or a long-horizon loop — not before.
  • I measured averages. For about two weeks my only latency number was a mean, which hid a p99 of six seconds behind a perfectly respectable 1.4.

And the thing I keep coming back to: the project that will look weakest on a portfolio — no interface, no demo, just tracing and a CI gate — is the one that found a bug I could not otherwise have found, and it is the one every engineer I have shown this to asked about first. Building is roughly a third of the work. Knowing whether the thing works, and why it stopped, is the rest.

If you are building any of this and want to compare notes — particularly on the citation gate, which I am still not fully happy with — the comments are open. Start with the production RAG guide for the retrieval theory behind project 1, and the AI engineering toolbox if you are still choosing tools.

#Performance #LLM #Agents #CI/CD #Agentic AI #RAG

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles

AI 13 min

Connecting OpenAI to a Real Backend: Lessons From SEO Automation

The naive version worked perfectly on fifty pages and failed four ways on twelve thousand — none of them about prompt quality. Wiring a model into a real backend is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle, and it needs everything batch processing always needed.

May 28, 2026
AI 15 min

The 940ms Voice Agent: Where Every Millisecond Goes

A conversation has a social latency budget, not a technical one — go past a second and people think the thing is broken. How a voice agent went from 1,530ms to 940ms to first audio, why every millisecond came from overlapping stages rather than speeding them up, the speculative optimisation that cost 38% more tokens for 60ms, and what a user hears when a dependency vanishes mid-sentence.

Apr 12, 2026
AI 14 min

Your AI System Is Failing Right Now and Nothing Is Alerting

On a Tuesday my RAG system got quietly worse. Every request returned 200, no alert fired, and the answers stayed fluent and confident while being subtly about the wrong thing. Why AI systems fail silently, what a trace must actually record, why the mean is a lie, and the forty-minute root cause that would have taken days.

Mar 27, 2026