Local Inference

← All models

Qwen3.8 27B Q4_K_M

Qwen3.8 27B · 27.3B dense · llama.cpp · GGUF Q4_K_M · 17.44 GB · bartowski/Qwen3.8-27B-GGUF @125a02a

Generation, empty context
31.8 t/s
± 0.24
Generation at 16k context
28.2 t/s
Prompt processing
949 t/s
Peak VRAM
16.9 GB
during llama-bench
GPU power while generating
246 W
Energy efficiency
0.13 tok/J

Settings

Hosted chat config. Its llama-bench numbers match 32k-q8kv (llama-bench takes no -c); it is benchmarked so the hosted config has a published run.

llama-server -m Qwen3.8-27B-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a qwen3.8-27b-q4_k_m-gguf/64k-q8kv --fit off -c 65536 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 --presence-penalty 0
SettingValue
ctx65536
gpu_layers999
flash_attnon
cache_type_kq8_0
cache_type_vq8_0
batch2048
ubatch512
threads6
n_cpu_moe0
parallel1
load_modemmap
cache_ram_mib0
spec
reasoning_format
reasoning_budget
extra_args["--temp","1.0","--top-p","0.95","--top-k","20","--min-p","0","--presence-penalty","0"]

Highlighted rows differ between this model's configs.

Speed

Generation vs context depth

All configs of this model; the selected one is solid.

Prompt processing vs context depth

TestDeptht/s± sdRepsGPU WTok/J
pp512 llama-bench094973.651516.29
pp512 llama-bench4k1,09125.152274.81
pp512 llama-bench16k98715.452404.11
tg128 llama-bench031.80.2452460.13
tg128 llama-bench4k30.90.2152450.13
tg128 llama-bench16k28.20.1252440.12

GPU power during the run

VRAM during the run

Quality

Not measured yet: perplexity and KL-divergence against a reference quant.

Task evals

Not measured yet.

Run history

DateKindEngine buildDriverDurationFlags
2026-09-15 18:47speedllama.cpp@91f6a6cf3595.99.022.8 min