Can a 3B Model Do the Job? Benchmarking Local LLMs on One Laptop
Three models, one laptop, thirty prompts, five runs each. Tokens per second, time to first token, memory and schema compliance for Llama 3.2 3B, Phi-4-mini and Mistral 7B — plus why I shipped the model that lost on speed, and what the Q4 to Q5 quantisation trade actually costs.
Every time I mention running a language model locally, someone asks the same reasonable question: why bother, when a frontier API is one HTTP call away? The honest answer is that "one HTTP call away" quietly assumes three things — a network, a budget, and permission to send the data — and in real work at least one of those is regularly missing.
So I stopped arguing about it and measured. Three models, one laptop, the same thirty prompts, five runs each, everything written to CSV. What came back changed which model I shipped, and for a reason I would not have guessed.
The case
Four constraints that make local the only option
None of these are hypothetical, and each one bites in a different place:
| Constraint | When it bites | What an API cannot fix |
|---|---|---|
| Privacy | Health records, legal documents, anything under a data-residency rule | A processing agreement is not the same as the data never leaving |
| Latency floor | Voice, autocomplete, anything a human waits on | The round trip is physics; you cannot optimise it away |
| Cost at scale | Classification or extraction on every row of a large table | Per-token pricing multiplied by millions of rows |
| No connectivity | Edge devices, factory floors, aircraft, rural deployments | There is no call to make |
There is a fifth reason that is less often stated and mattered most to me: a local model is a fallback that cannot go down with the same outage as your primary. When the cloud model in my voice agent times out, a 3B model on the same box answers in 260 ms. It is a worse answer. It is enormously better than silence.
Method
Measure before you form an opinion
The single most common mistake here is benchmarking by feel — running a few prompts, noticing one model "seems snappier", and shipping that. Two numbers matter and they move independently:
A summariser reads five thousand tokens and writes two hundred: it lives and dies on prefill. A chat assistant reads two hundred and writes five hundred: it lives on decode. The same model can be the better choice for one and the worse choice for the other.
def measure(model: str, prompt: str) -> Sample:
tracemalloc_peak = rss_watcher.start()
t0 = time.perf_counter()
first = None
tokens = 0
for chunk in ollama.stream(model=model, prompt=prompt, temperature=0.0):
if first is None:
first = time.perf_counter() # TTFT: the first token, not the first byte
tokens += 1
end = time.perf_counter()
return Sample(
model=model,
ttft_ms=(first - t0) * 1000,
decode_tps=tokens / (end - first), # NOT tokens / total — that hides prefill
total_ms=(end - t0) * 1000,
peak_rss_mb=rss_watcher.stop(),
)
tokens / total_time, which folds prefill into the decode rate and makes every model look slower on long prompts — by different amounts, depending on how fast its prefill is. The two phases have to be timed separately or the comparison is meaningless.
Results
Three models, one machine, thirty prompts
Ryzen 7 5800H, 32 GB RAM, RTX 3060 6 GB. Five runs per prompt, median reported. Same prompts, same order, same thermal state.
| Model | Quant | Memory | Cold load | TTFT | Tokens/sec | Valid JSON |
|---|---|---|---|---|---|---|
| Llama 3.2 3B | Q4_K_M | 2.4 GB | 1.9 s | 180 ms | 61.3 | 96.7% |
| Phi-4-mini 3.8B | Q4_K_M | 2.9 GB | 2.3 s | 210 ms | 52.8 | 98.3% |
| Mistral 7B | Q4_K_M | 4.6 GB | 3.8 s | 340 ms | 28.4 | 91.2% |
| Mistral 7B | Q5_K_M | 5.4 GB | 4.4 s | 390 ms | 22.1 | 93.8% |
Llama 3.2 is more than twice as fast as Mistral 7B and uses half the memory. If throughput were the question, the benchmark would end here.
The decision
The metric that actually chose the model
My use for a local model is structured extraction — take a messy sentence, return a validated JSON object. For that job, tokens per second is nearly irrelevant and schema compliance is everything, because an invalid object costs a full retry: another prefill, another decode, and a doubled latency for that request.
Run the arithmetic on 1,000 extractions:
| Model | Raw t/s | Invalid | Retries | Effective throughput |
|---|---|---|---|---|
| Llama 3.2 3B | 61.3 | 3.3% | 33 | 59.3 |
| Phi-4-mini | 52.8 | 1.7% | 17 | 51.9 |
| Mistral 7B Q4 | 28.4 | 8.8% | 88 | 26.1 |
Llama still wins on raw speed, and I shipped Phi-4-mini anyway — because in this system a failed extraction is not a slow result, it is a wrong one that reaches a user, and halving the failure rate was worth 14% of throughput I was not short of. That is a judgement about the workload, not about the models. The point is that I could only make it because both numbers existed.
Determinism
Constrain, validate, retry once
Raw schema violation rate across all three models was 8.4%. One retry with the validation error fed back into the prompt took that to 0.9%:
class Extraction(BaseModel):
intent: Literal["question", "command", "chitchat"]
entities: list[str]
confidence: float = Field(ge=0.0, le=1.0)
def extract(text: str, retries: int = 1) -> Extraction | None:
prompt = EXTRACT_PROMPT.format(text=text)
for attempt in range(retries + 1):
raw = ollama.generate(prompt, format="json", temperature=0.0)
try:
return Extraction.model_validate_json(raw)
except ValidationError as e:
if attempt == retries:
log.warning("extract.failed", text=text[:80], error=str(e))
return None # fail visibly — never silently
prompt = REPAIR_PROMPT.format(raw=raw, error=e)
retries: int = 1 is deliberate. Not three, not a while loop. A model that produced invalid output twice with the error handed straight back to it is not going to produce gold on the fifth attempt — it is going to produce a latency spike.
Temperature: the cheapest experiment I ran
Same thirty prompts, five runs each, two settings:
3 of 30 prompts produced any variation at all across five runs.
The three that did were open-ended questions where several phrasings are equally correct.
27 of 30 produced different output between runs.
Including six where the JSON structure itself changed — different key ordering, an optional field present in one run and absent in the next.
If you are extracting structured data and your temperature is not zero, you have chosen variance and received nothing for it. Keep the heat for the tasks where several answers are genuinely valid.
The trade everyone asks about
Quantisation: Q4 versus Q5
The received wisdom is that higher-bit quantisation is better quality for more memory. It is — the question is whether the exchange rate is any good, and that is machine- and task-specific enough that you have to measure it yourself.
22% of throughput and 800 MB, for 2.6 percentage points of schema compliance. I declined it. But notice what makes that a decision rather than a rule: on a machine with memory to spare running a batch job overnight, the same numbers point the other way, because nobody is waiting and 2.6 points is 26 fewer failures per thousand.
Honesty
What these numbers do not tell you
- One machine. Everything here is an RTX 3060 with 6 GB. A model that fits in VRAM behaves completely differently from one that spills to system RAM, and the 7B at Q5 was close to that edge. On a 24 GB card the whole table reorders.
- Thermals. A sustained ten-minute run dropped throughput about 8% as the laptop heated. Every number above is from a cool machine, which is the friendliest possible reading.
- Thirty prompts is a signal, not a verdict. It is enough to separate 98.3% from 91.2%. It is nowhere near enough to separate 98.3% from 97.9%.
- "Quality" here means schema compliance only. I did not measure whether the extracted entities were the right entities. That needs a golden set, which is a different piece of work.
- Model versions move. These are point-in-time results against specific builds. Pin your versions and re-run rather than trusting anyone's table, including this one.
The checklist
- Time prefill and decode separately — folding them together makes long-prompt comparisons meaningless.
- Report TTFT and tokens/sec, never a single "speed".
- Record peak memory, and note whether the model fit in VRAM.
- Same machine, same prompts, same thermal state, five runs, report the median.
- Measure the metric your workload actually cares about — for me that was schema compliance, not throughput.
- Compute effective throughput including retries, or you are comparing best cases.
- Temperature 0 for anything structured, and prove it with a variance run.
- Validate with a schema, retry exactly once, then fail visibly.
- Test the quantisation trade yourself rather than inheriting a rule of thumb.
- Write results to CSV and commit them, so the next run is a comparison.
Where to start on Monday
Install Ollama, pull a 3B model, and time thirty of your own prompts. Not thirty benchmark prompts — thirty of the things your system actually gets asked, because that distribution is the only one that matters.
The number that surprises you will not be the throughput. It will be how usable a 2.4 GB model is at the specific narrow job you needed doing, and how much of what you were paying for was generality you never used.
- Ollama
- TTFT
- decode t/s
- peak RSS
- Q4_K_M
- Pydantic
- retry-once
- effective throughput
This is project two of the four, and the model benchmarked here is the fallback that catches a cloud timeout in the voice agent. If you want the measurement discipline applied to answer quality rather than speed, that is the eval set post. Running different hardware and getting a different ordering? Post your numbers in the comments — that is exactly the data this table is missing.
Comments (0)
No comments yet
Be the first to share a thought on this article.
Join the conversation