Local Inference

← All models

Qwen3.8 27B Q4_K_M

Qwen3.8 27B · 27.3B dense · llama.cpp · GGUF Q4_K_M · 17.44 GB · bartowski/Qwen3.8-27B-GGUF @125a02a

Generation, empty context
31.7 t/s
± 0.29
Generation at 16k context
28.2 t/s
Prompt processing
1,037 t/s
Peak VRAM
16.9 GB
during llama-bench
GPU power while generating
247 W
Energy efficiency
0.13 tok/J

Settings

llama-server -m Qwen3.8-27B-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a qwen3.8-27b-q4_k_m-gguf/32k-q8kv --fit off -c 32768 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 --presence-penalty 0
SettingValue
ctx32768
gpu_layers999
flash_attnon
cache_type_kq8_0
cache_type_vq8_0
batch2048
ubatch512
threads6
n_cpu_moe0
parallel1
load_modemmap
cache_ram_mib0
spec
reasoning_format
reasoning_budget
extra_args["--temp","1.0","--top-p","0.95","--top-k","20","--min-p","0","--presence-penalty","0"]

Highlighted rows differ between this model's configs.

Speed

Generation vs context depth

All configs of this model; the selected one is solid.

Prompt processing vs context depth

TestDeptht/s± sdRepsGPU WTok/J
pp512 llama-bench01,03712751745.97
pp512 llama-bench4k1,09123.352274.80
pp512 llama-bench16k98814.852404.11
tg128 llama-bench031.70.2952470.13
tg128 llama-bench4k30.80.1952450.13
tg128 llama-bench16k28.20.1052450.12

GPU power during the run

VRAM during the run

Quality

Not measured yet: perplexity and KL-divergence against a reference quant.

Task evals

Not measured yet.

Run history

DateKindEngine buildDriverDurationFlags
2026-09-15 18:45speedllama.cpp@91f6a6cf3595.99.022.6 min