Qwen3.8 27B Q4_K_M
Generation, empty context
31.7 t/s
± 0.29
Generation at 16k context
28.2 t/s
Prompt processing
1,037 t/s
Peak VRAM
16.9 GB
during llama-bench
GPU power while generating
247 W
Energy efficiency
0.13 tok/J
Settings
llama-server -m Qwen3.8-27B-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a qwen3.8-27b-q4_k_m-gguf/32k-q8kv --fit off -c 32768 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 --presence-penalty 0
| Setting | Value |
|---|---|
ctx | 32768 |
gpu_layers | 999 |
flash_attn | on |
cache_type_k | q8_0 |
cache_type_v | q8_0 |
batch | 2048 |
ubatch | 512 |
threads | 6 |
n_cpu_moe | 0 |
parallel | 1 |
load_mode | mmap |
cache_ram_mib | 0 |
spec | — |
reasoning_format | — |
reasoning_budget | — |
extra_args | ["--temp","1.0","--top-p","0.95","--top-k","20","--min-p","0","--presence-penalty","0"] |
Highlighted rows differ between this model's configs.
Speed
Generation vs context depth
All configs of this model; the selected one is solid.
Prompt processing vs context depth
| Test | Depth | t/s | ± sd | Reps | GPU W | Tok/J |
|---|---|---|---|---|---|---|
| pp512 llama-bench | 0 | 1,037 | 127 | 5 | 174 | 5.97 |
| pp512 llama-bench | 4k | 1,091 | 23.3 | 5 | 227 | 4.80 |
| pp512 llama-bench | 16k | 988 | 14.8 | 5 | 240 | 4.11 |
| tg128 llama-bench | 0 | 31.7 | 0.29 | 5 | 247 | 0.13 |
| tg128 llama-bench | 4k | 30.8 | 0.19 | 5 | 245 | 0.13 |
| tg128 llama-bench | 16k | 28.2 | 0.10 | 5 | 245 | 0.12 |
GPU power during the run
VRAM during the run
Quality
Not measured yet: perplexity and KL-divergence against a reference quant.
Task evals
Not measured yet.
Run history
| Date | Kind | Engine build | Driver | Duration | Flags |
|---|---|---|---|---|---|
| 2026-09-15 18:45 | speed | llama.cpp@91f6a6cf3 | 595.99.02 | 2.6 min |