Gemma 3 4B IT Q4_K_M
Generation, empty context
180 t/s
± 0.83
Generation at 16k context
163 t/s
Prompt processing
8,196 t/s
Peak VRAM
3.6 GB
during llama-bench
GPU power while generating
207 W
Energy efficiency
0.87 tok/J
Settings
llama-server -m gemma-3-4b-it-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a gemma-3-4b-it-q4_k_m-gguf/32k-f16kv --fit off -c 32768 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk f16 -ctv f16 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics
| Setting | Value |
|---|---|
ctx | 32768 |
gpu_layers | 999 |
flash_attn | on |
cache_type_k | f16 |
cache_type_v | f16 |
batch | 2048 |
ubatch | 512 |
threads | 6 |
n_cpu_moe | 0 |
parallel | 1 |
load_mode | mmap |
cache_ram_mib | 0 |
spec | — |
reasoning_format | — |
reasoning_budget | — |
extra_args | [] |
Highlighted rows differ between this model's configs.
Speed
Generation vs context depth
All configs of this model; the selected one is solid.
Prompt processing vs context depth
| Test | Depth | t/s | ± sd | Reps | GPU W | Tok/J |
|---|---|---|---|---|---|---|
| pp512 llama-bench | 0 | 8,196 | 503 | 5 | 100 | 81.6 |
| pp512 llama-bench | 4k | 7,662 | 381 | 5 | 169 | 45.3 |
| pp512 llama-bench | 16k | 6,875 | 347 | 5 | 208 | 33.0 |
| tg128 llama-bench | 0 | 180 | 0.83 | 5 | 207 | 0.87 |
| tg128 llama-bench | 4k | 171 | 0.08 | 5 | 214 | 0.80 |
| tg128 llama-bench | 16k | 163 | 0.28 | 5 | 215 | 0.76 |
GPU power during the run
VRAM during the run
Quality
Not measured yet: perplexity and KL-divergence against a reference quant.
Task evals
Not measured yet.
Run history
| Date | Kind | Engine build | Driver | Duration | Flags |
|---|---|---|---|---|---|
| 2026-09-15 18:41 | speed | llama.cpp@91f6a6cf3 | 595.99.02 | 0.6 min |