Local Inference

Model benchmarks

Every config ranked for coding work on one RTX 3090. The score weights coding 50%, agents & tools 30% and reasoning 20%, and combines them as a weighted geometric mean, so being weak in any one area pulls it down. Answers must finish within a thinking allowance taken from each config’s own measured speed, so faster configs get more room to reason. How it's measured

No capability evals published yet. Every config gets the quick tier overnight; until then, the Speed tab has what each one delivers.