Suitable for short, straightforward conversations
The quickest eligible model at the balanced Q4 setting.
Tools
Estimate how fast a model generates tokens on your hardware, whether it fits in memory, and whether it suits your use case. Adjust any control and the estimate updates immediately.
24 GB · 936 GB/s · Ampere
Used $800–$1,000 · 26.7 GB VRAM/$1k · Read the hardware guide
Recommended for interactive chat: 12-30+ tok/s
Only models meeting this estimated decode speed are shown.
Three picks from the curated open-weight catalog that fit this system and speed filter. Select up to three to compare.
Suitable for short, straightforward conversations
The quickest eligible model at the balanced Q4 setting.
Stronger instruction following and reasoning
Balances model capability with responsive decode.
Slower, but with higher dense-model capacity
The largest model this system can run at a usable single-stream decode speed.
About these estimates
tok/s is estimated single-stream decode speed after prompt processing — not prefill speed or multi-user throughput.
Speed and memory start from published model size, KV-cache structure, device capacity and memory bandwidth, then apply architecture, runtime, offload and multi-device factors. The confidence label flags where those factors rely more heavily on interpolation.
Prices are indicative US whole-system street-price bands, updated July 2026; recommendation cards identify used-component mixes. Verify any shortlist on your own workload before buying.