Local Inference

Local Inference

Local Inference is my lab for running open-weight language models on a single RTX 3090. I tune each model's settings (context length, KV-cache precision, offload) and measure what the card actually delivers: generation and prompt speed as context fills up, peak VRAM, GPU power and tokens per joule. Every number comes from a scripted, hash-locked config, so results are reproducible.

How it's measured →

Fastest generation
180 t/s
Generation at 16k context
163 t/s
Best efficiency
0.87 tok/J
Benchmarked
2 models
5 configs

Benchmarks

Every model and config side by side: speed as the context fills, VRAM, power and efficiency, with the exact launch command for each.

View benchmarks →