Gemma 3 4B IT Q4_K_M
Generation, empty context
166 t/s
± 0.44
Generation at 16k context
156 t/s
Prompt processing
7,855 t/s
Peak VRAM
3.4 GB
during llama-bench
GPU power while generating
210 W
Energy efficiency
0.79 tok/J
Settings
llama-server -m gemma-3-4b-it-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a gemma-3-4b-it-q4_k_m-gguf/32k-q8kv --fit off -c 32768 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics
| Setting | Value |
|---|---|
ctx | 32768 |
gpu_layers | 999 |
flash_attn | on |
cache_type_k | q8_0 |
cache_type_v | q8_0 |
batch | 2048 |
ubatch | 512 |
threads | 6 |
n_cpu_moe | 0 |
parallel | 1 |
load_mode | mmap |
cache_ram_mib | 0 |
spec | — |
reasoning_format | — |
reasoning_budget | — |
extra_args | [] |
Highlighted rows differ between this model's configs.
Speed
Generation vs context depth
All configs of this model; the selected one is solid.
Prompt processing vs context depth
| Test | Depth | t/s | ± sd | Reps | GPU W | Tok/J |
|---|---|---|---|---|---|---|
| pp512 llama-bench | 0 | 7,855 | 574 | 5 | — | — |
| pp512 llama-bench | 4k | 7,369 | 426 | 5 | 196 | 37.7 |
| pp512 llama-bench | 16k | 6,598 | 332 | 5 | 211 | 31.2 |
| tg128 llama-bench | 0 | 166 | 0.44 | 5 | 210 | 0.79 |
| tg128 llama-bench | 4k | 161 | 0.66 | 5 | 214 | 0.75 |
| tg128 llama-bench | 16k | 156 | 0.60 | 5 | 222 | 0.70 |
GPU power during the run
VRAM during the run
Quality
Not measured yet: perplexity and KL-divergence against a reference quant.
Task evals
Not measured yet.
Run history
| Date | Kind | Engine build | Driver | Duration | Flags |
|---|---|---|---|---|---|
| 2026-09-15 18:42 | speed | llama.cpp@91f6a6cf3 | 595.99.02 | 0.6 min |