Skip to content
A AhsanLab.Tech
AI 13 min read · March 11, 2026

Can a 3B Model Do the Job? Benchmarking Local LLMs on One Laptop

Three models, one laptop, thirty prompts, five runs each. Tokens per second, time to first token, memory and schema compliance for Llama 3.2 3B, Phi-4-mini and Mistral 7B — plus why I shipped the model that lost on speed, and what the Q4 to Q5 quantisation trade actually costs.

A Ahsan Habib Save

Every time I mention running a language model locally, someone asks the same reasonable question: why bother, when a frontier API is one HTTP call away? The honest answer is that "one HTTP call away" quietly assumes three things — a network, a budget, and permission to send the data — and in real work at least one of those is regularly missing.

So I stopped arguing about it and measured. Three models, one laptop, the same thirty prompts, five runs each, everything written to CSV. What came back changed which model I shipped, and for a reason I would not have guessed.

61.3tokens / sec, best caseLlama 3.2 3B, Q4
98.3%valid JSON, best casea different model
−22%throughput, Q4 → Q5for +2.6pp quality

The case

Four constraints that make local the only option

None of these are hypothetical, and each one bites in a different place:

ConstraintWhen it bitesWhat an API cannot fix
Privacy Health records, legal documents, anything under a data-residency rule A processing agreement is not the same as the data never leaving
Latency floor Voice, autocomplete, anything a human waits on The round trip is physics; you cannot optimise it away
Cost at scale Classification or extraction on every row of a large table Per-token pricing multiplied by millions of rows
No connectivity Edge devices, factory floors, aircraft, rural deployments There is no call to make

There is a fifth reason that is less often stated and mattered most to me: a local model is a fallback that cannot go down with the same outage as your primary. When the cloud model in my voice agent times out, a 3B model on the same box answers in 260 ms. It is a worse answer. It is enormously better than silence.

Method

Measure before you form an opinion

The single most common mistake here is benchmarking by feel — running a few prompts, noticing one model "seems snappier", and shipping that. Two numbers matter and they move independently:

WHERE THE TIME GOES IN ONE LOCAL GENERATION MODEL LOAD once — 1.9 to 3.8 s PREFILL reads the prompt DECODE one token at a time DONE the last token TIME TO FIRST TOKEN what a person perceives as "did it hear me" scales with PROMPT length TOKENS / SECOND how fast the answer arrives after that scales with ANSWER length a long prompt with a short answer is a TTFT problem · a short prompt with a long answer is a throughput problem
FIGURE 1 — TWO NUMBERS, NOT ONE · Averaging them into "speed" is how you pick the wrong model for your workload.

A summariser reads five thousand tokens and writes two hundred: it lives and dies on prefill. A chat assistant reads two hundred and writes five hundred: it lives on decode. The same model can be the better choice for one and the worse choice for the other.

bench/run.py — CSV, because screenshots are not data
def measure(model: str, prompt: str) -> Sample:
    tracemalloc_peak = rss_watcher.start()
    t0 = time.perf_counter()
    first = None
    tokens = 0

    for chunk in ollama.stream(model=model, prompt=prompt, temperature=0.0):
        if first is None:
            first = time.perf_counter()      # TTFT: the first token, not the first byte
        tokens += 1

    end = time.perf_counter()
    return Sample(
        model=model,
        ttft_ms=(first - t0) * 1000,
        decode_tps=tokens / (end - first),   # NOT tokens / total — that hides prefill
        total_ms=(end - t0) * 1000,
        peak_rss_mb=rss_watcher.stop(),
    )
The measurement bug I shipped first My first harness computed throughput as tokens / total_time, which folds prefill into the decode rate and makes every model look slower on long prompts — by different amounts, depending on how fast its prefill is. The two phases have to be timed separately or the comparison is meaningless.

Results

Three models, one machine, thirty prompts

Ryzen 7 5800H, 32 GB RAM, RTX 3060 6 GB. Five runs per prompt, median reported. Same prompts, same order, same thermal state.

ModelQuantMemoryCold loadTTFTTokens/secValid JSON
Llama 3.2 3BQ4_K_M2.4 GB1.9 s180 ms61.396.7%
Phi-4-mini 3.8BQ4_K_M2.9 GB2.3 s210 ms52.898.3%
Mistral 7BQ4_K_M4.6 GB3.8 s340 ms28.491.2%
Mistral 7BQ5_K_M5.4 GB4.4 s390 ms22.193.8%
Llama 3.2 3B61.3 t/s
Phi-4-mini52.8 t/s
Mistral 7B Q428.4 t/s
Mistral 7B Q522.1 t/s

Llama 3.2 is more than twice as fast as Mistral 7B and uses half the memory. If throughput were the question, the benchmark would end here.

The decision

The metric that actually chose the model

My use for a local model is structured extraction — take a messy sentence, return a validated JSON object. For that job, tokens per second is nearly irrelevant and schema compliance is everything, because an invalid object costs a full retry: another prefill, another decode, and a doubled latency for that request.

Run the arithmetic on 1,000 extractions:

ModelRaw t/sInvalidRetriesEffective throughput
Llama 3.2 3B61.33.3%3359.3
Phi-4-mini52.81.7%1751.9
Mistral 7B Q428.48.8%8826.1

Llama still wins on raw speed, and I shipped Phi-4-mini anyway — because in this system a failed extraction is not a slow result, it is a wrong one that reaches a user, and halving the failure rate was worth 14% of throughput I was not short of. That is a judgement about the workload, not about the models. The point is that I could only make it because both numbers existed.

"Fastest" is not a property of a model. It is a property of a model, a workload, and a machine — and the benchmark that omits any of the three is measuring somebody else's decision.

Determinism

Constrain, validate, retry once

Raw schema violation rate across all three models was 8.4%. One retry with the validation error fed back into the prompt took that to 0.9%:

local/structured.py
class Extraction(BaseModel):
    intent: Literal["question", "command", "chitchat"]
    entities: list[str]
    confidence: float = Field(ge=0.0, le=1.0)

def extract(text: str, retries: int = 1) -> Extraction | None:
    prompt = EXTRACT_PROMPT.format(text=text)
    for attempt in range(retries + 1):
        raw = ollama.generate(prompt, format="json", temperature=0.0)
        try:
            return Extraction.model_validate_json(raw)
        except ValidationError as e:
            if attempt == retries:
                log.warning("extract.failed", text=text[:80], error=str(e))
                return None              # fail visibly — never silently
            prompt = REPAIR_PROMPT.format(raw=raw, error=e)

retries: int = 1 is deliberate. Not three, not a while loop. A model that produced invalid output twice with the error handed straight back to it is not going to produce gold on the fifth attempt — it is going to produce a latency spike.

Temperature: the cheapest experiment I ran

Same thirty prompts, five runs each, two settings:

temperature 0.0

3 of 30 prompts produced any variation at all across five runs.

The three that did were open-ended questions where several phrasings are equally correct.

temperature 0.7

27 of 30 produced different output between runs.

Including six where the JSON structure itself changed — different key ordering, an optional field present in one run and absent in the next.

If you are extracting structured data and your temperature is not zero, you have chosen variance and received nothing for it. Keep the heat for the tasks where several answers are genuinely valid.

The trade everyone asks about

Quantisation: Q4 versus Q5

The received wisdom is that higher-bit quantisation is better quality for more memory. It is — the question is whether the exchange rate is any good, and that is machine- and task-specific enough that you have to measure it yourself.

MISTRAL 7B · WHAT Q5 COSTS AND WHAT IT BUYS Throughput 28.4 t/s 22.1 t/s · −22% Memory 4.6 GB 5.4 GB · +800 MB Valid JSON 91.2% 93.8% · +2.6pp Q4_K_M Q5_K_M bars are scaled within each row
FIGURE 2 — THE EXCHANGE RATE · Two rows of cost against one row of benefit. On this machine, for this task, that is a trade worth declining.

22% of throughput and 800 MB, for 2.6 percentage points of schema compliance. I declined it. But notice what makes that a decision rather than a rule: on a machine with memory to spare running a batch job overnight, the same numbers point the other way, because nobody is waiting and 2.6 points is 26 fewer failures per thousand.

Honesty

What these numbers do not tell you

  • One machine. Everything here is an RTX 3060 with 6 GB. A model that fits in VRAM behaves completely differently from one that spills to system RAM, and the 7B at Q5 was close to that edge. On a 24 GB card the whole table reorders.
  • Thermals. A sustained ten-minute run dropped throughput about 8% as the laptop heated. Every number above is from a cool machine, which is the friendliest possible reading.
  • Thirty prompts is a signal, not a verdict. It is enough to separate 98.3% from 91.2%. It is nowhere near enough to separate 98.3% from 97.9%.
  • "Quality" here means schema compliance only. I did not measure whether the extracted entities were the right entities. That needs a golden set, which is a different piece of work.
  • Model versions move. These are point-in-time results against specific builds. Pin your versions and re-run rather than trusting anyone's table, including this one.

The checklist

  • Time prefill and decode separately — folding them together makes long-prompt comparisons meaningless.
  • Report TTFT and tokens/sec, never a single "speed".
  • Record peak memory, and note whether the model fit in VRAM.
  • Same machine, same prompts, same thermal state, five runs, report the median.
  • Measure the metric your workload actually cares about — for me that was schema compliance, not throughput.
  • Compute effective throughput including retries, or you are comparing best cases.
  • Temperature 0 for anything structured, and prove it with a variance run.
  • Validate with a schema, retry exactly once, then fail visibly.
  • Test the quantisation trade yourself rather than inheriting a rule of thumb.
  • Write results to CSV and commit them, so the next run is a comparison.

Where to start on Monday

Install Ollama, pull a 3B model, and time thirty of your own prompts. Not thirty benchmark prompts — thirty of the things your system actually gets asked, because that distribution is the only one that matters.

The number that surprises you will not be the throughput. It will be how usable a 2.4 GB model is at the specific narrow job you needed doing, and how much of what you were paying for was generality you never used.

  • Ollama
  • TTFT
  • decode t/s
  • peak RSS
  • Q4_K_M
  • Pydantic
  • retry-once
  • effective throughput

This is project two of the four, and the model benchmarked here is the fallback that catches a cloud timeout in the voice agent. If you want the measurement discipline applied to answer quality rather than speed, that is the eval set post. Running different hardware and getting a different ordering? Post your numbers in the comments — that is exactly the data this table is missing.

#Performance #AI Tools #LLM #Agents #Prompt Engineering

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles

AI 13 min

Connecting OpenAI to a Real Backend: Lessons From SEO Automation

The naive version worked perfectly on fifty pages and failed four ways on twelve thousand — none of them about prompt quality. Wiring a model into a real backend is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle, and it needs everything batch processing always needed.

May 28, 2026
AI 15 min

The 940ms Voice Agent: Where Every Millisecond Goes

A conversation has a social latency budget, not a technical one — go past a second and people think the thing is broken. How a voice agent went from 1,530ms to 940ms to first audio, why every millisecond came from overlapping stages rather than speeding them up, the speculative optimisation that cost 38% more tokens for 60ms, and what a user hears when a dependency vanishes mid-sentence.

Apr 12, 2026
AI 14 min

Your AI System Is Failing Right Now and Nothing Is Alerting

On a Tuesday my RAG system got quietly worse. Every request returned 200, no alert fired, and the answers stayed fluent and confident while being subtly about the wrong thing. Why AI systems fail silently, what a trace must actually record, why the mean is a lie, and the forty-minute root cause that would have taken days.

Mar 27, 2026