Qwen3.8 27B Q4_K_M
Generation, empty context
32.4 t/s
± 0.20
Generation at 16k context
30.3 t/s
Prompt processing
965 t/s
Peak VRAM
17.3 GB
during llama-bench
GPU power while generating
245 W
Energy efficiency
0.13 tok/J
Settings
llama-server -m Qwen3.8-27B-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a qwen3.8-27b-q4_k_m-gguf/32k-f16kv --fit off -c 32768 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk f16 -ctv f16 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 --presence-penalty 0
| Setting | Value |
|---|---|
ctx | 32768 |
gpu_layers | 999 |
flash_attn | on |
cache_type_k | f16 |
cache_type_v | f16 |
batch | 2048 |
ubatch | 512 |
threads | 6 |
n_cpu_moe | 0 |
parallel | 1 |
load_mode | mmap |
cache_ram_mib | 0 |
spec | — |
reasoning_format | — |
reasoning_budget | — |
extra_args | ["--temp","1.0","--top-p","0.95","--top-k","20","--min-p","0","--presence-penalty","0"] |
Highlighted rows differ between this model's configs.
Speed
Generation vs context depth
All configs of this model; the selected one is solid.
Prompt processing vs context depth
| Test | Depth | t/s | ± sd | Reps | GPU W | Tok/J |
|---|---|---|---|---|---|---|
| pp512 llama-bench | 0 | 965 | 71.7 | 5 | 173 | 5.58 |
| pp512 llama-bench | 4k | 1,109 | 26.2 | 5 | 233 | 4.75 |
| pp512 llama-bench | 16k | 998 | 15.8 | 5 | 237 | 4.21 |
| tg128 llama-bench | 0 | 32.4 | 0.20 | 5 | 245 | 0.13 |
| tg128 llama-bench | 4k | 31.7 | 0.10 | 5 | 244 | 0.13 |
| tg128 llama-bench | 16k | 30.3 | 0.16 | 5 | 242 | 0.13 |
GPU power during the run
VRAM during the run
Quality
Not measured yet: perplexity and KL-divergence against a reference quant.
Task evals
Not measured yet.
Run history
| Date | Kind | Engine build | Driver | Duration | Flags |
|---|---|---|---|---|---|
| 2026-09-15 18:42 | speed | llama.cpp@91f6a6cf3 | 595.99.02 | 2.3 min |