Local Inference

← All models

Gemma 3 4B IT Q4_K_M

Gemma 3 4B IT · 3.9B dense · llama.cpp · GGUF Q4_K_M · 2.49 GB · ggml-org/gemma-3-4b-it-GGUF @d097622

Generation, empty context
166 t/s
± 0.44
Generation at 16k context
156 t/s
Prompt processing
7,855 t/s
Peak VRAM
3.4 GB
during llama-bench
GPU power while generating
210 W
Energy efficiency
0.79 tok/J

Settings

llama-server -m gemma-3-4b-it-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a gemma-3-4b-it-q4_k_m-gguf/32k-q8kv --fit off -c 32768 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics
SettingValue
ctx32768
gpu_layers999
flash_attnon
cache_type_kq8_0
cache_type_vq8_0
batch2048
ubatch512
threads6
n_cpu_moe0
parallel1
load_modemmap
cache_ram_mib0
spec
reasoning_format
reasoning_budget
extra_args[]

Highlighted rows differ between this model's configs.

Speed

Generation vs context depth

All configs of this model; the selected one is solid.

Prompt processing vs context depth

TestDeptht/s± sdRepsGPU WTok/J
pp512 llama-bench07,8555745
pp512 llama-bench4k7,369426519637.7
pp512 llama-bench16k6,598332521131.2
tg128 llama-bench01660.4452100.79
tg128 llama-bench4k1610.6652140.75
tg128 llama-bench16k1560.6052220.70

GPU power during the run

VRAM during the run

Quality

Not measured yet: perplexity and KL-divergence against a reference quant.

Task evals

Not measured yet.

Run history

DateKindEngine buildDriverDurationFlags
2026-09-15 18:42speedllama.cpp@91f6a6cf3595.99.020.6 min