Most GPU buying advice is written for gamers, and it is close to useless for AI work. Frame rates measure something a language model never does. The three specifications that decide whether a card is any good for local AI — memory capacity, memory bandwidth, and framework support — barely appear in a typical review.
This guide is organised around those three, in that order.
Our picks at a glance
-
Best Overall
NVIDIA GeForce RTX 5090
The only consumer card with enough memory for 32B-class models plus long context
-
Best Value
NVIDIA GeForce RTX 5070 Ti
16 GB and 896 GB/s at 300 W — the sensible middle of the range
-
Best Budget
NVIDIA GeForce RTX 5060 Ti 16GB
The cheapest route to 16 GB of CUDA memory
-
Best High-End Option
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
96 GB on one card, for 70B-class work without splitting layers
-
Most VRAM per Dollar
AMD Radeon RX 7900 XTX
24 GB for text generation at a lower price than NVIDIA charges
-
Best Premium
NVIDIA GeForce RTX 4090
Still excellent, and often the best capability per dollar on the used market
How we chose
Our full approach is in the review methodology. For graphics cards specifically, we weight in this order:
- Memory capacity, because it determines what you can run at all. This is a threshold, not a gradient: a card that is 2 GB short of your model is the wrong card, however fast it is.
- Memory bandwidth, because once a model fits, token generation is almost entirely bandwidth-bound. A useful ceiling is bandwidth ÷ model size in memory.
- Framework support, because a theoretical capability you cannot reach through your toolchain is not a capability.
- Power and thermals, because a 575 W card imposes real costs on the rest of the build.
Compute throughput matters too, but it only becomes the binding constraint for image generation and training. For the text generation most people are doing, it is rarely what limits you.
The specifications side by side
| Product | VRAM | Bandwidth | Power | Runs up to | Local LLMs | Where to buy |
|---|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition NVIDIA | 96 GB | 1792 GB/s | 600 W (Max-Q variant is configurable to 300 W) | 70B at Q8, or 120B-class at Q4, entirely in VRAM | Excellent | Check Price on Amazon NVIDIA RTX PRO 6000 Blackwell Workstation Edition at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5090 NVIDIA | 32 GB | 1792 GB/s | 575 W total board power; 1,000 W system PSU recommended | 32B at Q4 entirely in VRAM, with room for long context | Excellent | Check Price on Amazon NVIDIA GeForce RTX 5090 at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 4090 NVIDIA | 24 GB | 1008 GB/s | 450 W total board power; 850 W system PSU recommended | 32B at Q4 with modest context, or 14B at Q8 | Excellent | Check Price on Amazon NVIDIA GeForce RTX 4090 at Amazon — opens in a new tab |
| AMD Radeon RX 7900 XTX AMD | 24 GB | 960 GB/s | 355 W total board power; 800 W system PSU recommended | 32B at Q4 under llama.cpp or Ollama | Good | Check Price on Amazon AMD Radeon RX 7900 XTX at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5080 NVIDIA | 16 GB | 960 GB/s | 360 W total board power; 850 W system PSU recommended | 14B at Q4 with long context, or 32B at Q4 with partial CPU offload | Good | Check Price on Amazon NVIDIA GeForce RTX 5080 at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5070 Ti NVIDIA | 16 GB | 896 GB/s | 300 W total board power; 750 W system PSU recommended | 14B at Q4 with comfortable context headroom | Good | Check Price on Amazon NVIDIA GeForce RTX 5070 Ti at Amazon — opens in a new tab |
| AMD Radeon RX 9070 XT AMD | 16 GB | 645 GB/s | 304 W total board power; 750 W system PSU recommended | 14B at Q4 | Workable with caveats | Check Price on Amazon AMD Radeon RX 9070 XT at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5060 Ti 16GB NVIDIA | 16 GB | 448 GB/s | 180 W total board power; 600 W system PSU recommended | 14B at Q4, at roughly half the token rate of a 5070 Ti | Workable with caveats | Check Price on Amazon NVIDIA GeForce RTX 5060 Ti 16GB at Amazon — opens in a new tab |
The recommendations
Best overall
Best Overall
NVIDIA GeForce RTX 5090
4.1/5AI Gear Stack review score: 4.1 out of 5
Best for The fastest single-card local inference you can buy without moving to workstation pricing
The 32 GB frame buffer and 1.8 TB/s of bandwidth make this the only consumer card that runs 32B-class models comfortably in VRAM. The cost is real: 575 W, a 1,000 W PSU and a case that can move the heat.
- VRAM
- 32 GB
- Bandwidth
- 1792 GB/s
- GPU
- GB202 Blackwell, 21,760 CUDA cores
- Runs up to
- 32B at Q4 entirely in VRAM, with room for long context
Strengths
- 32 GB VRAM clears the 30B-class barrier that 24 GB cards sit just under
- 1,792 GB/s is roughly 78% more bandwidth than an RTX 4090, and bandwidth is what sets token speed
- Native FP4 support meaningfully increases effective capacity on supported runtimes
- CUDA remains the path of least resistance for every local AI toolchain
Trade-offs
- 575 W board power demands a serious PSU and case airflow
- Still short of the 48 GB needed to hold a 70B model at Q4
- Physically large; verify clearance before buying
- Availability and street pricing have been erratic since launch
The RTX 5090 earns this position on one number: 32 GB.
That figure clears a specific and important threshold. A 32-billion-parameter model at 4-bit quantisation needs roughly 19 GB for weights, plus several gigabytes of KV cache at a useful context length. On a 24 GB card that fits, but only just, and you spend your time trimming context. On 32 GB it fits with room to work.
The bandwidth is the second reason. At 1,792 GB/s the 5090 has roughly 78% more memory bandwidth than an RTX 4090, and since generation speed scales with bandwidth, that translates fairly directly into tokens per second on any model that fits both cards.
What it does not do is reach 70B. A 70B model at 4-bit needs about 42 GB of weights before context, and no amount of optimisation closes a 10 GB gap. If 70B is your requirement, this is not your card.
The costs are real and worth stating plainly. 575 W of board power needs a 1,000 W-class supply, a case that can move the heat, and a tolerance for the noise that comes with removing it.
Best value
Best Value
NVIDIA GeForce RTX 5070 Ti
3.9/5AI Gear Stack review score: 3.9 out of 5
Best for The cheapest sensible entry to 16 GB of fast NVIDIA VRAM
Within about 7% of a 5080 on memory bandwidth with the same 16 GB, at meaningfully lower power and price. For most people choosing between the two, this is the better buy.
- VRAM
- 16 GB
- Bandwidth
- 896 GB/s
- GPU
- GB203 Blackwell, 8,960 CUDA cores
- Runs up to
- 14B at Q4 with comfortable context headroom
Strengths
- Same 16 GB and near-identical bandwidth to the RTX 5080
- 300 W fits comfortably in an existing mid-range build
- Strong price-to-capability ratio for image generation
Trade-offs
- 16 GB caps language-model ambitions at roughly 14B
- No capability advantage over the 5080 — only a price one
The RTX 5070 Ti is the card we recommend most often, because the specification that matters is nearly identical to the RTX 5080’s.
Both have 16 GB on a 256-bit bus. The 5080 has 896 GB/s versus… 960 GB/s. That is a 7% bandwidth advantage for a meaningful price premium, on the specification that most determines AI performance. The 5080’s extra shader count matters for gaming and for image generation; for language models it largely does not.
At 300 W it also drops into an existing mid-range build without a power supply upgrade, which the 5080 and 5090 frequently do not.
The limitation is the same 16 GB ceiling. That comfortably covers 14B-class models — which is where a great deal of genuinely useful local AI happens — and does not reach 32B.
Best budget
Best Budget
NVIDIA GeForce RTX 5060 Ti 16GB
3.6/5AI Gear Stack review score: 3.6 out of 5
Best for Getting 16 GB of CUDA VRAM into a small or low-power machine
The cheapest way to get 16 GB of CUDA memory. The 128-bit bus means it loads big models it then runs slowly — capacity without the bandwidth to exploit it.
- VRAM
- 16 GB
- Bandwidth
- 448 GB/s
- GPU
- GB206 Blackwell, 4,608 CUDA cores
- Runs up to
- 14B at Q4, at roughly half the token rate of a 5070 Ti
Strengths
- 16 GB at the lowest tier NVIDIA offers it in
- 180 W runs happily on a modest PSU and in a small case
- Strong fit for an always-on inference box
Trade-offs
- 448 GB/s is a genuine bottleneck: roughly half a 5070 Ti
- The 8 GB variant shares the name and is a different product — check carefully
If the goal is simply to get 16 GB of CUDA memory into a machine for the least money, this is the card.
It comes with an honest caveat that most coverage glosses over: the 128-bit bus gives it 448 GB/s, roughly half the RTX 5070 Ti. It will load the same models and run them at roughly half the speed. That is capacity without the bandwidth to exploit it.
For an always-on inference box, a homelab node, or a machine where 180 W and a small case matter more than tokens per second, that trade is sensible. If you are going to sit and watch it generate, it is not.
Be careful with the name. There is an 8 GB variant sold under the same model number. For AI work they are different products, and the 8 GB card is not one we would recommend.
Best high-end
Best High-End Option
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
4.2/5AI Gear Stack review score: 4.2 out of 5
Best for Running or fine-tuning 70B-class models on a single card
96 GB of ECC GDDR7 at RTX 5090 bandwidth. It is the single-card answer to 70B-class work, priced as a professional tool rather than a consumer component.
- VRAM
- 96 GB
- Bandwidth
- 1792 GB/s
- GPU
- GB202 Blackwell, 24,064 CUDA cores
- Runs up to
- 70B at Q8, or 120B-class at Q4, entirely in VRAM
Strengths
- 96 GB holds a 70B model at Q8 with context to spare
- Same 1,792 GB/s bandwidth as the RTX 5090
- ECC memory and the professional driver branch for long unattended runs
- Max-Q variant fits 300 W power envelopes
Trade-offs
- Professional pricing, several times an RTX 5090
- 600 W in the full-power configuration
- Overkill for anything that fits in 32 GB
This is the single-card answer to 70B, and the first genuine one.
96 GB of ECC GDDR7 at the same 1,792 GB/s as an RTX 5090 means a 70B model at 8-bit fits entirely in memory, with context to spare — no layer splitting, no PCIe bottleneck between cards, no multi-GPU runtime quirks. For fine-tuning, where memory pressure is far higher than inference, the headroom matters even more.
It is priced as a professional tool, several times an RTX 5090. The Max-Q variant, configurable down to 300 W, is worth knowing about if you are building into a constrained power or thermal envelope.
Best memory per dollar
Most VRAM per Dollar
AMD Radeon RX 7900 XTX
3.9/5AI Gear Stack review score: 3.9 out of 5
Best for 24 GB of VRAM for text generation, at a lower price than the NVIDIA equivalent
24 GB and 960 GB/s for less money than a comparable NVIDIA card. The hardware is not the problem — the software ecosystem is, and how much that costs you depends entirely on your workload.
- VRAM
- 24 GB
- Bandwidth
- 960 GB/s
- GPU
- Navi 31 RDNA 3, 96 compute units
- Runs up to
- 32B at Q4 under llama.cpp or Ollama
Strengths
- 24 GB at a considerably lower price than NVIDIA charges
- 960 GB/s matches an RTX 5080 despite the older architecture
- Text generation through llama.cpp and Ollama is well supported and fast
Trade-offs
- ROCm support is narrower and more fragile than CUDA
- Image generation and fine-tuning workflows frequently assume CUDA
- Expect to spend time on setup that an NVIDIA card would not require
24 GB and 960 GB/s, typically for meaningfully less than the NVIDIA cards that match it on memory.
For text generation specifically, this is a strong proposition. llama.cpp and Ollama both have solid ROCm and Vulkan back-ends, and a 32B model at 4-bit runs well.
The honest caveat is everything else. Image generation pipelines, fine-tuning frameworks and most research code assume CUDA. Some of it works on ROCm, some works after effort, and some does not work at all. If your workload is “run language models” the hardware is excellent value. If it is “try whatever appears on GitHub this month”, the friction is real and recurring.
Still excellent
Best Premium
NVIDIA GeForce RTX 4090
4/5AI Gear Stack review score: 4 out of 5
Best for A 24 GB CUDA card at second-hand prices
Still an excellent local inference card. 24 GB and 1 TB/s remain competitive, and on the used market it is often the best capability-per-dollar available.
- VRAM
- 24 GB
- Bandwidth
- 1008 GB/s
- GPU
- AD102 Ada Lovelace, 16,384 CUDA cores
- Runs up to
- 32B at Q4 with modest context, or 14B at Q8
Strengths
- 24 GB handles 32B at Q4 and most image workflows
- 1,008 GB/s is within 45% of an RTX 5090
- Mature, universally supported in every local AI toolchain
Trade-offs
- No native FP4 support
- No longer in production; supply is second-hand
- 450 W with the same 12VHPWR connector caveats
The RTX 4090 remains a genuinely good local AI card. 24 GB runs 32B models at 4-bit, and 1,008 GB/s is within striking distance of current hardware.
It is out of production, so buying means the second-hand market. When prices are sensible it is frequently the best capability-per-dollar available. Verify the card’s history and inspect the power connector.
What actually determines performance
Capacity is a cliff, not a slope
The most common expensive mistake is buying a card that is almost big enough.
When a model does not fit, the runtime splits it — some layers on the GPU, the rest on the CPU — and the CPU portion runs across the PCIe bus at a fraction of the speed. Performance does not degrade gracefully. It falls off a cliff, typically by five to fifteen times.
This is why we recommend sizing for the model you actually want to run, with headroom, rather than buying the fastest card in budget and hoping.
Bandwidth sets the ceiling
Generating each token requires reading the model’s active weights out of memory. That makes generation memory-bandwidth-bound, and gives a ceiling you can calculate before buying:
theoretical tokens/second ≈ memory bandwidth ÷ model size in memory
A 19 GB model on an RTX 5090’s 1,792 GB/s ceilings out around 94 tokens/second; expect 60–75 in practice. The same model on an RTX 5060 Ti’s 448 GB/s ceilings around 23. Same model, same capacity, four times the speed difference — entirely explained by the bus width.
Compute matters for images, not text
Image generation inverts the hierarchy. Diffusion models are small enough that capacity is rarely the constraint, and the work is compute-bound rather than bandwidth-bound.
If Stable Diffusion or ComfyUI is your main workload, weight shader count and tensor throughput far more heavily, and 16 GB is generally sufficient.
What to look for
Buy the memory tier your workload needs, then optimise within it. Decide whether you are a 16, 24, 32 or 48 GB user first. Comparing a 16 GB card against a 24 GB card on price is comparing two different capabilities.
Check the bus width, not just the capacity. 16 GB on 256-bit and 16 GB on 128-bit are the same capacity and roughly a factor of two apart in speed.
Budget the whole system. A 575 W card needs a 1,000 W supply, adequate case airflow, and physical clearance. These are not optional extras.
Leave the PCIe slot argument alone. For single-GPU inference, PCIe generation and lane count barely matter — weights are transferred once. It becomes significant only when splitting a model across cards.
Consider the previous generation seriously. In AI work, memory capacity ages far more slowly than compute. A 24 GB card from two generations back can beat a current 16 GB card at the thing you actually want to do.
If none of these fit
If your requirement is 70B-class models and a workstation GPU is out of budget, a graphics card is probably the wrong purchase entirely.
A unified-memory machine — the Framework Desktop with 128 GB, a Mac Studio, or NVIDIA’s DGX Spark — will run models no consumer card can load, at a fraction of the power. They are much slower per token, but slow is infinitely faster than does-not-run. We cover the trade in Best AI Workstations.
Common questions
Is the RTX 5090 worth it over the RTX 4090 for AI?
For language models, yes — 32 GB versus 24 GB is the difference between running a 32B model with long context and running it with a short one, and the 5090 has about 78% more bandwidth. For image generation, where 24 GB is rarely the limit, the gap is much smaller and a used 4090 is often better value.
Can I use two GPUs to get more VRAM?
Yes. Runtimes such as llama.cpp, Ollama and vLLM split model layers across cards, so two 24 GB cards give you 48 GB and will run a 70B model at 4-bit. Throughput is less than double a single card because activations cross the PCIe bus at each layer boundary, and you need the power delivery and cooling for both.
Does AMD work for local AI?
For text generation through llama.cpp or Ollama, yes, and a 24 GB Radeon is genuinely good value. For image generation, fine-tuning, or anything that assumes CUDA, expect real friction. Our NVIDIA vs AMD comparison goes through it properly.
How much VRAM do I actually need?
It depends entirely on model size and context length. As a planning figure, a 4-bit model needs roughly 0.6 GB per billion parameters plus about 20% for context and overhead. The full arithmetic is in How Much VRAM Do You Need for Local AI?
Is a used RTX 3090 still a good option?
It remains one of the better value propositions in local AI: 24 GB of VRAM at around 936 GB/s, usually for considerably less than a current card. The trade-offs are age, no warranty, higher idle power, and no support for newer numeric formats. If budget is the binding constraint and you can verify the card’s history, it is a reasonable choice.
Do I need a top-tier CPU to go with the GPU?
No. Once model weights are resident in GPU memory, the CPU handles tokenisation and orchestration and little else. Any modern mid-range processor is fine. Spend the difference on VRAM.
Continue your research
- How Much VRAM Do You Need for Local AI? — size your requirement before you spend
- RTX 5090 vs RTX 5080 for AI — the two cards most people cross-shop
- NVIDIA vs AMD for Local AI — what the ecosystem gap actually costs
- Best GPUs for Ollama — the same hardware, assessed for one specific runtime