Local Inference

Model benchmarks

Local model performance, measured on one machine. How it's measured

One row per model, using its fastest config (generation speed, empty context). Speeds are means of repeated llama-bench runs on this machine.

ModelEngineSizeConfigCtxPP t/sTG t/s ↓TG t/s deepVRAMTok/J
Gemma 3 4B IT Q4_K_M
Q4_K_M · dense
llama.cpp2.49 GB32k ctx, f16 KV32k8,196180163 @16k3.6 GB0.87
Qwen3.8 27B Q4_K_M
Q4_K_M · dense
llama.cpp17.44 GB32k ctx, f16 KV32k96532.430.3 @16k17.3 GB0.13

Generation speed

Tokens per second at empty context. Click a bar to open the model.

Speed vs weight size

Bigger weights mean more memory traffic per token.