I Built Four AI Engineering Projects in 12 Weeks — Here Is Every Number
A production RAG system with enforced citations, a local model benchmark, an observability layer, and a real-time voice agent — built in sequence as one system, measured end to end. Why I deliberately skipped fine-tuning, the numbers from every phase, and the four things I got wrong first.
Twelve weeks ago I had the same problem every engineer moving into AI has: a folder full of tutorial chatbots that all worked perfectly on the three questions I had thought to ask them, and no idea whether any of it resembled the work an actual AI team does.
So I picked four projects, gave myself twelve weeks, and made one rule: every claim has to be a number I measured myself. No "significantly faster." No "much more accurate." A number, from a script anyone can re-run.
This is the write-up. Four projects, what each one taught me, the numbers that came out of them, and the three days I lost to a bug I would never have found without the boring project nobody builds.
First: what I deliberately did not build
Most "AI portfolio" lists put fine-tuning in the middle — LoRA, QLoRA, DPO, training curves. I dropped it on purpose, and I want to explain why before anything else, because the reasoning shaped every other decision here.
There is a line between two jobs that get advertised under one name:
An AI engineer composes models into reliable systems. An ML engineer changes the model's weights.
Retrieval, tool calling, constrained generation, evaluation, observability, streaming, failure handling — that is composition, and that is the job I want. Fine-tuning is a genuinely valuable skill, but it is a different one, and a portfolio that gestures weakly at both says less than one that is unmistakably deep in one.
There is also an honest engineering argument. Fine-tuning is what you reach for after you have demonstrated a ceiling — after you can show, with a golden evaluation set and real metrics, that careful prompting and better retrieval cannot get you there. I did not have that evidence twelve weeks ago. Nobody starting out does. Building the measurement apparatus first is not a detour around fine-tuning; it is the thing that tells you whether you need it.
The four projects are one system
I did not build four unrelated repos. Each project is a layer of the same thing, and the order matters: project 1 produces the evaluation set that project 3 gates on, and project 2 produces the local model that project 4 falls back to when the cloud times out.
Project 1 — A RAG system that refuses to answer
The corpus is Laravel 12, Filament and Livewire documentation: 1,240 pages, 8,900 chunks. I picked it because I work in it daily, which means I can judge whether an answer is right. That sounds obvious and it is the single most important decision in the project — a golden evaluation set built on a corpus you cannot personally grade is a set of guesses in a CSV.
Phase 1 — the forty-line version
Chunk at 500–800 tokens with roughly 100 tokens of overlap, embed, store in Chroma, retrieve top-5, stuff into the prompt, ask for citations. It works, it is genuinely satisfying, and it is where most portfolio projects stop.
Here is what it scored on my evaluation set:
| Configuration | Faithfulness | Context recall | Refusal correctness | p95 latency |
|---|---|---|---|---|
| Dense vector only, top-5 | 0.71 | 0.68 | 12% | 1.9 s |
| + BM25 hybrid, RRF fusion | 0.79 | 0.84 | 18% | 2.1 s |
| + cross-encoder rerank | 0.88 | 0.89 | 24% | 3.4 s |
| + citation enforcement gate | 0.94 | 0.89 | 91% | 3.5 s |
Read the refusal correctness column, not the faithfulness one. That is the percentage of the 38 deliberately unanswerable questions in my set where the system correctly said "the documentation does not cover this" instead of inventing something. The naive pipeline got that right 12% of the time. It confidently made up an answer to almost nine out of ten questions it had no business answering — and every one of those answers was well-formatted, plausible, and wrong.
That number is the whole reason this project exists.
Phase 2 — hybrid retrieval and the rerank
withoutGlobalScopes; the vector search catches "how do I stop a scope applying." Reciprocal Rank Fusion needs no score normalisation, which is why it beat the weighted-sum blend I tried first.The fusion step is the part people over-engineer. I started with a weighted sum of normalised scores and spent a day tuning the weight, then threw it away for Reciprocal Rank Fusion, which ignores scores entirely and only looks at rank position:
# retrieval/fuse.py — RRF needs no score normalisation, which is the point.
# BM25 scores and cosine similarities live on incomparable scales; ranks don't.
def reciprocal_rank_fusion(rankings: list[list[str]], k: int = 60) -> list[str]:
scores: dict[str, float] = {}
for ranking in rankings:
for position, chunk_id in enumerate(ranking):
scores[chunk_id] = scores.get(chunk_id, 0.0) + 1.0 / (k + position + 1)
return sorted(scores, key=scores.get, reverse=True)
candidates = reciprocal_rank_fusion([
bm25.search(query, top_k=20),
vectors.search(embed(query), top_k=20),
])[:30]
# The reranker is the expensive step, so it runs once, on 30 candidates, not 8,900.
top_chunks = reranker.rank(query, candidates)[:6]
The reranker costs 380 ms at p95 and moved faithfulness from 0.79 to 0.88. That is the best latency-for-quality trade in the entire system, and it is one model call.
The citation gate
This is the piece I would keep if I could keep only one. Before an answer is returned, every sentence in it is checked against the retrieved chunks. If a claim is not supported, the answer does not go out.
# answer/gate.py — a refusal is a feature, not an error path.
result = llm.generate(prompt, chunks=top_chunks)
unsupported = [
claim for claim in split_claims(result.text)
if not any(supports(chunk, claim) for chunk in top_chunks)
]
if unsupported:
log.info("gate.refused", claims=len(unsupported), query=query)
return Refusal(
message="The documentation I have does not cover this.",
nearest=top_chunks[:2], # show what we DID find — never a dead end
)
return Answer(text=result.text, citations=result.citations)
Refusal correctness went from 24% to 91%. The remaining 9% are questions where the docs contain something adjacent enough that the check passes — a partial answer to a question about a related feature. I have not solved those and I am not going to pretend otherwise.
Phase 3 — the golden set and the CI gate
180 question/answer pairs, hand-verified, 38 of them deliberately unanswerable. It took two full days to build and it is the most valuable artefact in all four projects. Ragas runs it on every pull request; if faithfulness drops below 0.85 or context recall below 0.80, the build fails.
In six weeks that gate blocked three of my own pull requests. One was a chunking change I was certain was an improvement. It was not.
Project 2 — Running it all offline on a laptop
The question I kept getting asked: why bother with a 3B model when a frontier API is one HTTP call away? The answer is that "one HTTP call away" assumes a network, a budget, and permission to send the data — and in real work at least one of those three is regularly missing. Privacy rules, latency floors, cost at scale, and edge deployments with no connectivity are not edge cases; they are Tuesday.
Ollama, a FastAPI wrapper, and a benchmark harness that writes CSV rather than screenshots. Everything below is on one machine — Ryzen 7 5800H, 32 GB RAM, RTX 3060 6 GB — over the same 30-prompt suite, five runs each.
| Model | Quant | Memory | Time to first token | Tokens/sec | Valid JSON |
|---|---|---|---|---|---|
| Llama 3.2 3B | Q4_K_M | 2.4 GB | 180 ms | 61.3 | 96.7% |
| Phi-4-mini 3.8B | Q4_K_M | 2.9 GB | 210 ms | 52.8 | 98.3% |
| Mistral 7B | Q4_K_M | 4.6 GB | 340 ms | 28.4 | 91.2% |
| Mistral 7B | Q5_K_M | 5.4 GB | 390 ms | 22.1 | 93.8% |
The Q4 → Q5 row is the one worth staring at. Going from Q4_K_M to Q5_K_M cost 22% of throughput and 800 MB of memory and bought 2.6 percentage points of schema compliance. On this hardware, for this task, that is a bad trade — and I would not have known that from any blog post, because it depends entirely on the task and the machine.
Phi-4-mini won on the metric I actually cared about, despite losing on the one everyone quotes. That is the entire lesson of the project.
Structure, validation, retry
Raw schema violation rate was 8.4%. One retry with the validation error fed back into the prompt took it to 0.9%:
# local/structured.py — constrain, validate, retry once, then fail honestly.
class Extraction(BaseModel):
intent: Literal["question", "command", "chitchat"]
entities: list[str]
confidence: float = Field(ge=0.0, le=1.0)
def extract(text: str, retries: int = 1) -> Extraction | None:
prompt = EXTRACT_PROMPT.format(text=text)
for attempt in range(retries + 1):
raw = ollama.generate(prompt, format="json", temperature=0.0)
try:
return Extraction.model_validate_json(raw)
except ValidationError as e:
if attempt == retries:
log.warning("extract.failed", text=text[:80], error=str(e))
return None # fail visibly, never silently
prompt = REPAIR_PROMPT.format(raw=raw, error=e)
Note retries: int = 1. Not 3, not a while True. A model that produced garbage twice with the error handed back to it is not going to produce gold on the fifth attempt — it is going to produce a latency spike and a bill.
The temperature study was the cheapest experiment here and the one that changed how I think. Same 30 prompts, five runs each: at temperature=0, 3 of 30 prompts produced any variation at all across runs. At 0.7, 27 of 30 did. If you are extracting structured data and your temperature is not zero, you have chosen variance for no reason.
Project 3 — The boring one that found the bug
This is the project that looks worst in a portfolio and taught me the most. No pretty frontend. No demo you can show someone in ten seconds. Just tracing on every step of the RAG pipeline, four metrics tracked over time, and a CI gate.
I self-hosted Langfuse rather than using a managed tier, purely so I would not hit a usage ceiling while iterating. Every request records which chunks were retrieved, how the reranker reordered them, the exact prompt sent, the response, and the token counts.
| Metric | Value | Why this one |
|---|---|---|
| Latency p50 | 1.4 s | What a typical user feels |
| Latency p95 | 3.4 s | What your angriest user feels |
| Latency p99 | 6.1 s | Where the reranker queues up |
| Cost per request | $0.0031 | Multiply by traffic before promising anything |
| Citation coverage | 94.2% | Proportion of answers fully grounded |
| Failure rate | 1.8% | 0.6% errors, 1.2% refusals — tracked separately |
Note the mean is nowhere on that table. The average of 1.4 s and 6.1 s is a number that describes nobody's experience.
The Tuesday that justified the whole project
Six weeks in, answer quality got noticeably worse. Not broken — worse. Nothing errored, nothing alerted, the system returned confident answers all day. If I had been running the naive setup I would have assumed I was imagining it.
Instead I opened the dashboard, found context recall had dropped from 0.89 to 0.61, and used the traces to see that the retrieved chunks had gone subtly off-topic. Root cause: an unpinned dependency had bumped the embedding model version. The new vectors and the ones already in the index were computed by different models, so every similarity score in the store was quietly meaningless.
Forty minutes from noticing to root cause, because the traces existed. Without them I would have spent days rewriting prompts that were never the problem — and I know that, because that is exactly what I did for three days earlier in the project, before this layer existed.
# .github/workflows/eval.yml — a prompt change is a behaviour change.
- name: RAG evaluation
run: python -m eval.run --dataset golden/v3.jsonl --out results.json
- name: Gate on regression
run: |
python -m eval.gate results.json \
--min-faithfulness 0.85 \
--min-context-recall 0.80
# exits non-zero → the build fails → the change does not merge
Prompts live in prompts/*.yaml, versioned next to the code and diffed in review. A one-word prompt edit can move faithfulness three points. Treating that as a config change rather than a code change is how teams ship regressions they cannot explain.
Project 4 — Voice, where every millisecond is visible
I picked the voice track over the vision and log-analysis options because a conversation is unforgiving about latency in a way a dashboard never is. Nobody notices 400 ms in an API. Everybody notices it in a reply.
Mic → ASR → retrieval → LLM → TTS → speaker, orchestrated over WebSockets with a structured event protocol. Phase one was just getting audio to flow end to end, which took longer than I expected and optimised nothing.
Phase two was the interesting part: decomposing the latency instead of measuring it as one number.
Total time to first audio started at 1,530 ms and finished at 940 ms (p95: 1.8 s). Three changes did nearly all of it:
- Start retrieving on partial transcripts rather than waiting for ASR to finalise. Overlapping those two stages saved ~260 ms and cost some wasted retrievals, which is a trade I will take every time.
- Chunk TTS by sentence. Speaking the first sentence while the model is still generating the second removes the entire tail of generation from perceived latency.
- Keep the reranker warm. A cold cross-encoder added 600 ms to the first request after any idle period, which is exactly the request a person is most likely to be judging.
Phase 3 — what happens when things break
This is where project 2 stopped being a separate repo and became a component. Every dependency gets a timeout and a documented fallback:
| Failure | Detected by | Degradation | Cost |
|---|---|---|---|
| ASR unavailable | WebSocket close / 3 s timeout | Switch to typed input, say so out loud | Voice input lost, session survives |
| Cloud LLM timeout | 4 s deadline | Fall back to local Llama 3.2 3B | +260 ms, shorter answers |
| TTS unavailable | 5 s timeout | Return the answer as text | No audio, answer still delivered |
| Retrieval empty | Zero chunks over threshold | Refuse, offer to rephrase | No answer — correctly |
The rule underneath all four rows: the system never hangs silently. It degrades visibly and says what it is doing. A voice assistant that goes quiet for eight seconds has failed more completely than one that says "let me try that a simpler way."
Replay mode was the last thing I built and the thing I would build first next time. 74 recorded sessions can be fed back through the pipeline deterministically, which turns "it did something weird once yesterday" from a shrug into a test case.
What I would tell someone starting this on Monday
Four other things I got wrong first:
- I optimised prompts for three days when the problem was retrieval. The model was doing fine with what it was handed; what it was handed was wrong. Without traces I could not see that, so I tuned the one thing I could see.
- I picked a corpus I could not grade. My first attempt used a research-paper set in a field I do not know. I could not tell a good answer from a confident one, which made the whole evaluation set worthless. Starting over on documentation I use daily was the right call.
- I reached for a framework too early. The first agent loop was LangGraph wrapping a single model call with two tools. The model SDK's own tool-runner did the same job in a third of the code. Frameworks earn their place when you have several coordinating agents or a long-horizon loop — not before.
- I measured averages. For about two weeks my only latency number was a mean, which hid a p99 of six seconds behind a perfectly respectable 1.4.
And the thing I keep coming back to: the project that will look weakest on a portfolio — no interface, no demo, just tracing and a CI gate — is the one that found a bug I could not otherwise have found, and it is the one every engineer I have shown this to asked about first. Building is roughly a third of the work. Knowing whether the thing works, and why it stopped, is the rest.
If you are building any of this and want to compare notes — particularly on the citation gate, which I am still not fully happy with — the comments are open. Start with the production RAG guide for the retrieval theory behind project 1, and the AI engineering toolbox if you are still choosing tools.
Comments (0)
No comments yet
Be the first to share a thought on this article.
Join the conversation