Model benchmarks
One row per model, using its fastest config (generation speed, empty context). Speeds are means of repeated llama-bench runs on this machine.
| Model | Engine | Size | Config | Ctx | PP t/s | TG t/s ↓ | TG t/s deep | VRAM | Tok/J |
|---|---|---|---|---|---|---|---|---|---|
| Gemma 3 4B IT Q4_K_M Q4_K_M · dense | llama.cpp | 2.49 GB | 32k ctx, f16 KV | 32k | 8,196 | 180 | 163 @16k | 3.6 GB | 0.87 |
| Qwen3.8 27B Q4_K_M Q4_K_M · dense | llama.cpp | 17.44 GB | 32k ctx, f16 KV | 32k | 965 | 32.4 | 30.3 @16k | 17.3 GB | 0.13 |
Generation speed
Tokens per second at empty context. Click a bar to open the model.
Speed vs weight size
Bigger weights mean more memory traffic per token.