Qwen3.8 27B Q4_K_M
Generation, empty context
31.8 t/s
± 0.24
Generation at 16k context
28.2 t/s
Prompt processing
949 t/s
Peak VRAM
16.9 GB
during llama-bench
GPU power while generating
246 W
Energy efficiency
0.13 tok/J
Settings
Hosted chat config. Its llama-bench numbers match 32k-q8kv (llama-bench takes no -c); it is benchmarked so the hosted config has a published run.
llama-server -m Qwen3.8-27B-Q4_K_M.gguf --host 127.0.0.1 --port 8080 -a qwen3.8-27b-q4_k_m-gguf/64k-q8kv --fit off -c 65536 -np 1 --cache-ram 0 -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 -t 6 -ncmoe 0 -lm mmap --metrics --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 --presence-penalty 0
| Setting | Value |
|---|---|
ctx | 65536 |
gpu_layers | 999 |
flash_attn | on |
cache_type_k | q8_0 |
cache_type_v | q8_0 |
batch | 2048 |
ubatch | 512 |
threads | 6 |
n_cpu_moe | 0 |
parallel | 1 |
load_mode | mmap |
cache_ram_mib | 0 |
spec | — |
reasoning_format | — |
reasoning_budget | — |
extra_args | ["--temp","1.0","--top-p","0.95","--top-k","20","--min-p","0","--presence-penalty","0"] |
Highlighted rows differ between this model's configs.
Speed
Generation vs context depth
All configs of this model; the selected one is solid.
Prompt processing vs context depth
| Test | Depth | t/s | ± sd | Reps | GPU W | Tok/J |
|---|---|---|---|---|---|---|
| pp512 llama-bench | 0 | 949 | 73.6 | 5 | 151 | 6.29 |
| pp512 llama-bench | 4k | 1,091 | 25.1 | 5 | 227 | 4.81 |
| pp512 llama-bench | 16k | 987 | 15.4 | 5 | 240 | 4.11 |
| tg128 llama-bench | 0 | 31.8 | 0.24 | 5 | 246 | 0.13 |
| tg128 llama-bench | 4k | 30.9 | 0.21 | 5 | 245 | 0.13 |
| tg128 llama-bench | 16k | 28.2 | 0.12 | 5 | 244 | 0.12 |
GPU power during the run
VRAM during the run
Quality
Not measured yet: perplexity and KL-divergence against a reference quant.
Task evals
Not measured yet.
Run history
| Date | Kind | Engine build | Driver | Duration | Flags |
|---|---|---|---|---|---|
| 2026-09-15 18:47 | speed | llama.cpp@91f6a6cf3 | 595.99.02 | 2.8 min |