Tools

Local AI Calculator

Estimate how fast a model generates tokens on your hardware, whether it fits in memory, and whether it suits your use case. Adjust any control and the estimate updates immediately.

HardwareArchitecture names such as Ampere identify a GPU generation. Memory capacity determines what fits; memory bandwidth strongly influences decode speed.

24 GB · 936 GB/s · Ampere

Used $800–$1,000 · 26.7 GB VRAM/$1k · Read the hardware guide

Primary use
1510203050100200400

Recommended for interactive chat: 12-30+ tok/s

Only models meeting this estimated decode speed are shown.

What to run on RTX 3090

Three picks from the curated open-weight catalog that fit this system and speed filter. Select up to three to compare.

Fastest at Q4

Gemma 4 E2B on Hugging Face

Q4_K_M234–294 tok/sExpected: 263 tok/s

Suitable for short, straightforward conversations

The quickest eligible model at the balanced Q4 setting.

Best balance

Qwen3.6 35B-A3B on Hugging Face

Q4_K_M118–149 tok/sExpected: 133 tok/s

Stronger instruction following and reasoning

Balances model capability with responsive decode.

Largest viable

Qwen3 32B on Hugging Face

Q4_K_M29–37 tok/sExpected: 33 tok/s

Slower, but with higher dense-model capacity

The largest model this system can run at a usable single-stream decode speed.

About these estimates

tok/s is estimated single-stream decode speed after prompt processing — not prefill speed or multi-user throughput.

Speed and memory start from published model size, KV-cache structure, device capacity and memory bandwidth, then apply architecture, runtime, offload and multi-device factors. The confidence label flags where those factors rely more heavily on interpolation.

Prices are indicative US whole-system street-price bands, updated July 2026; recommendation cards identify used-component mixes. Verify any shortlist on your own workload before buying.