Chat
Talk to the model this machine is hosting right now, running entirely on the RTX 3090.
Starting · Qwen3.8 27B Q4_K_M · 64k ctx, q8_0 KVPrivate: sign-in required.
Open chat ↗Local Inference is my lab for running open-weight language models on a single RTX 3090. I tune each model's settings (context length, KV-cache precision, offload) and measure what the card actually delivers: generation and prompt speed as context fills up, peak VRAM, GPU power and tokens per joule. Every number comes from a scripted, hash-locked config, so results are reproducible.
Talk to the model this machine is hosting right now, running entirely on the RTX 3090.
Starting · Qwen3.8 27B Q4_K_M · 64k ctx, q8_0 KVPrivate: sign-in required.
Open chat ↗Every model and config side by side: speed as the context fills, VRAM, power and efficiency, with the exact launch command for each.
View benchmarks →