Local Inference

← All models

Gemma 3 4B IT Q4_K_M

Gemma 3 4B IT · 3.9B dense · llama.cpp · GGUF Q4_K_M · 2.49 GB · ggml-org/gemma-3-4b-it-GGUF @d097622

Generation, empty context
180 t/s
± 0.83
Generation at 16k context
163 t/s
Prompt processing
8,196 t/s
Peak VRAM
3.6 GB
during llama-bench
GPU power while generating
207 W
Energy efficiency
0.87 tok/J

Settings

llama-server -m gemma-3-4b-it-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a gemma-3-4b-it-q4_k_m-gguf/32k-f16kv --fit off -c 32768 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk f16 -ctv f16 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics
SettingValue
ctx32768
gpu_layers999
flash_attnon
cache_type_kf16
cache_type_vf16
batch2048
ubatch512
threads6
n_cpu_moe0
parallel1
load_modemmap
cache_ram_mib0
spec
reasoning_format
reasoning_budget
extra_args[]

Highlighted rows differ between this model's configs.

Speed

Generation vs context depth

All configs of this model; the selected one is solid.

Prompt processing vs context depth

TestDeptht/s± sdRepsGPU WTok/J
pp512 llama-bench08,196503510081.6
pp512 llama-bench4k7,662381516945.3
pp512 llama-bench16k6,875347520833.0
tg128 llama-bench01800.8352070.87
tg128 llama-bench4k1710.0852140.80
tg128 llama-bench16k1630.2852150.76

GPU power during the run

VRAM during the run

Quality

Not measured yet: perplexity and KL-divergence against a reference quant.

Task evals

Not measured yet.

Run history

DateKindEngine buildDriverDurationFlags
2026-09-15 18:41speedllama.cpp@91f6a6cf3595.99.020.6 min