Local Inference

← All models

Qwen3.8 27B Q4_K_M

Qwen3.8 27B · 27.3B dense · llama.cpp · GGUF Q4_K_M · 17.44 GB · bartowski/Qwen3.8-27B-GGUF @125a02a

Generation, empty context
32.4 t/s
± 0.20
Generation at 16k context
30.3 t/s
Prompt processing
965 t/s
Peak VRAM
17.3 GB
during llama-bench
GPU power while generating
245 W
Energy efficiency
0.13 tok/J

Settings

llama-server -m Qwen3.8-27B-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a qwen3.8-27b-q4_k_m-gguf/32k-f16kv --fit off -c 32768 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk f16 -ctv f16 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 --presence-penalty 0
SettingValue
ctx32768
gpu_layers999
flash_attnon
cache_type_kf16
cache_type_vf16
batch2048
ubatch512
threads6
n_cpu_moe0
parallel1
load_modemmap
cache_ram_mib0
spec
reasoning_format
reasoning_budget
extra_args["--temp","1.0","--top-p","0.95","--top-k","20","--min-p","0","--presence-penalty","0"]

Highlighted rows differ between this model's configs.

Speed

Generation vs context depth

All configs of this model; the selected one is solid.

Prompt processing vs context depth

TestDeptht/s± sdRepsGPU WTok/J
pp512 llama-bench096571.751735.58
pp512 llama-bench4k1,10926.252334.75
pp512 llama-bench16k99815.852374.21
tg128 llama-bench032.40.2052450.13
tg128 llama-bench4k31.70.1052440.13
tg128 llama-bench16k30.30.1652420.13

GPU power during the run

VRAM during the run

Quality

Not measured yet: perplexity and KL-divergence against a reference quant.

Task evals

Not measured yet.

Run history

DateKindEngine buildDriverDurationFlags
2026-09-15 18:42speedllama.cpp@91f6a6cf3595.99.022.3 min