Almost everything this site tells you about buying a GPU for AI is wrong for image generation.
For language models the guidance is memory first, bandwidth second, compute a distant third. Diffusion inverts all of it: compute is what you are buying, and VRAM is rarely the thing that stops you. A card we would call limited for local LLMs can be close to ideal here.
Understanding why makes the whole category easy to shop.
Our picks at a glance
-
Best Overall
NVIDIA GeForce RTX 5080
Twice the shader throughput of the 5070 Ti, and 16 GB is plenty here
-
Best Value
NVIDIA GeForce RTX 5070 Ti
The same memory at lower cost and power, and perceptibly slower
-
Best for Image Generation
NVIDIA GeForce RTX 5090
Flux at full precision, video generation, entire chains resident
-
Best Budget
NVIDIA GeForce RTX 5060 Ti 16GB
16 GB cheaply — at roughly half the images per minute
-
Most VRAM per Dollar
NVIDIA GeForce RTX 3090
24 GB and 10,496 cores, on the used market
Why image generation is a different problem
The difference comes down to what the hardware actually does in each case.
A language model generating text reads its entire set of active weights out of memory to produce a single token, then does it again for the next one. The arithmetic per token is small; the reading is enormous. That makes it memory-bandwidth-bound, which is why VRAM capacity and bus width dominate our other guides.
A diffusion model generating an image loads a comparatively small network once, then runs twenty to fifty denoising steps over a latent tensor. Each step is dense arithmetic on data already sitting in fast memory. The reading is modest; the arithmetic is enormous. That makes it compute-bound.
The practical consequence: for images, the number that predicts your speed is the shader and tensor throughput, not the memory bandwidth. Doubling the CUDA cores roughly halves the time per image. Doubling the memory bandwidth does very little.
How much VRAM you actually need
Less than the language-model numbers on this site would suggest, until you hit the specific cases where it isn’t.
| Model family | Approximate VRAM at fp16 | Reduced precision |
|---|---|---|
| Stable Diffusion 1.5 | ~2 GB | — |
| SDXL | ~7 GB | ~5 GB |
| Flux.1 (12B) | ~24 GB | ~12 GB at fp8, ~8 GB quantised |
| Video generation models | Substantially more | Varies widely |
For SD 1.5 and SDXL, 8 GB is workable and 12 GB is comfortable. The large modern image models are the reason 16 GB has become the sensible floor, and Flux at full precision is the reason 24 GB is worth having.
Where VRAM does become the constraint
Four situations, and they are the ones that generate the “out of memory” posts:
- High resolution. Attention memory grows faster than linearly with the number of latent tokens, so generating at 2048 rather than 1024 costs far more than four times the memory. This is the most common cause of a sudden failure.
- Batch size. Generating four images at once costs roughly four times the working memory. It is also considerably faster per image than four separate runs, which is why people do it.
- ComfyUI workflows holding several models. A base model, a refiner, a ControlNet and an upscaler chained together is four models. ComfyUI will unload between stages, and keeping them resident is much faster.
- Video generation. The case where VRAM genuinely binds, and where 24 GB or more stops being optional.
The VAE decode trap
Worth knowing because it catches people at the worst moment.
The final step of generation decodes the latent back into pixels, and at high resolution that decode is a memory spike considerably larger than the generation that preceded it. People run a 30-step generation successfully and then run out of memory on the last operation.
Every mainstream interface offers tiled VAE decoding, which processes the image in sections. It is slightly slower and it removes the ceiling entirely. Turn it on before you need it.
The specifications
| Product | GPU | VRAM | Power | Image generation | Where to buy |
|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 NVIDIA | GB202 Blackwell, 21,760 CUDA cores | 32 GB | 575 W total board power; 1,000 W system PSU recommended | Excellent | Check Price on Amazon NVIDIA GeForce RTX 5090 at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5080 NVIDIA | GB203 Blackwell, 10,752 CUDA cores | 16 GB | 360 W total board power; 850 W system PSU recommended | Excellent | Check Price on Amazon NVIDIA GeForce RTX 5080 at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 3090 NVIDIA | GA102 Ampere, 10,496 CUDA cores | 24 GB | 350 W total board power; 850 W system PSU recommended | Good | Check Price on Amazon NVIDIA GeForce RTX 3090 at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5070 Ti NVIDIA | GB203 Blackwell, 8,960 CUDA cores | 16 GB | 300 W total board power; 750 W system PSU recommended | Excellent | Check Price on Amazon NVIDIA GeForce RTX 5070 Ti at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5060 Ti 16GB NVIDIA | GB206 Blackwell, 4,608 CUDA cores | 16 GB | 180 W total board power; 600 W system PSU recommended | Good | Check Price on Amazon NVIDIA GeForce RTX 5060 Ti 16GB at Amazon — opens in a new tab |
The recommendations
Best overall
Best Overall
NVIDIA GeForce RTX 5080
3.6/5AI Gear Stack review score: 3.6 out of 5
Best for Image generation and 14B-class language models, at half the power of a 5090
A fast card held back for language work by 16 GB of VRAM. For Stable Diffusion and ComfyUI it is close to ideal; for LLMs it sits on the wrong side of the 30B threshold.
- VRAM
- 16 GB
- Bandwidth
- 960 GB/s
- GPU
- GB203 Blackwell, 10,752 CUDA cores
- Runs up to
- 14B at Q4 with long context, or 32B at Q4 with partial CPU offload
Strengths
- 360 W is far easier to cool and power than the 5090
- 960 GB/s is ample for anything that fits in 16 GB
- Excellent for image generation, where 16 GB is rarely the constraint
- Substantially cheaper than the 5090 for two thirds of the shader count
Trade-offs
- 16 GB is the defining limitation for language models
- Half the memory of a 5090 for well over half the price
- Offloading a 32B model to system RAM costs most of the speed advantage
This is the card that looks mediocre in our language-model guides and excellent here, and the reason is the inversion described above.
10,752 CUDA cores is roughly twice the shader throughput of the 5070 Ti, and for compute-bound diffusion work that translates almost directly into images per minute. The 16 GB that limits it to 14B-class language models is, for image generation, ample — enough for SDXL at large batch sizes and for Flux at reduced precision with room for a ControlNet.
At 360 W it is also a far easier card to build around than the 5090, and it does not need a 1,000 W supply.
If your machine is primarily for image generation, this is the sensible top of the range.
Best value
Best Value
NVIDIA GeForce RTX 5070 Ti
3.9/5AI Gear Stack review score: 3.9 out of 5
Best for The cheapest sensible entry to 16 GB of fast NVIDIA VRAM
Within about 7% of a 5080 on memory bandwidth with the same 16 GB, at meaningfully lower power and price. For most people choosing between the two, this is the better buy.
- VRAM
- 16 GB
- Bandwidth
- 896 GB/s
- GPU
- GB203 Blackwell, 8,960 CUDA cores
- Runs up to
- 14B at Q4 with comfortable context headroom
Strengths
- Same 16 GB and near-identical bandwidth to the RTX 5080
- 300 W fits comfortably in an existing mid-range build
- Strong price-to-capability ratio for image generation
Trade-offs
- 16 GB caps language-model ambitions at roughly 14B
- No capability advantage over the 5080 — only a price one
8,960 CUDA cores and the same 16 GB, at 300 W and meaningfully less money.
For language models we recommend this over the 5080 because the two are nearly identical on the specification that matters. For image generation that argument does not hold — the 5080 genuinely is faster here, because compute is what you are buying.
It remains the better value, and it will generate images perceptibly more slowly. Which matters depends on whether you generate a few images a day or run batches for an hour.
If you run Flux at full precision, or generate video
Best for Image Generation
NVIDIA GeForce RTX 5090
4.1/5AI Gear Stack review score: 4.1 out of 5
Best for The fastest single-card local inference you can buy without moving to workstation pricing
The 32 GB frame buffer and 1.8 TB/s of bandwidth make this the only consumer card that runs 32B-class models comfortably in VRAM. The cost is real: 575 W, a 1,000 W PSU and a case that can move the heat.
- VRAM
- 32 GB
- Bandwidth
- 1792 GB/s
- GPU
- GB202 Blackwell, 21,760 CUDA cores
- Runs up to
- 32B at Q4 entirely in VRAM, with room for long context
Strengths
- 32 GB VRAM clears the 30B-class barrier that 24 GB cards sit just under
- 1,792 GB/s is roughly 78% more bandwidth than an RTX 4090, and bandwidth is what sets token speed
- Native FP4 support meaningfully increases effective capacity on supported runtimes
- CUDA remains the path of least resistance for every local AI toolchain
Trade-offs
- 575 W board power demands a serious PSU and case airflow
- Still short of the 48 GB needed to hold a 70B model at Q4
- Physically large; verify clearance before buying
- Availability and street pricing have been erratic since launch
The one case where the language-model logic and the image-generation logic agree.
32 GB holds Flux at full fp16 precision without quantisation compromises, keeps an entire ComfyUI chain resident between stages, and gives video generation the headroom it needs. 21,760 CUDA cores make it the fastest consumer card per image by a wide margin.
For SDXL work it is substantial overkill — you are paying for memory the workload never touches. For the largest current models it is the card that removes every constraint at once.
Best budget
Best Budget
NVIDIA GeForce RTX 5060 Ti 16GB
3.6/5AI Gear Stack review score: 3.6 out of 5
Best for Getting 16 GB of CUDA VRAM into a small or low-power machine
The cheapest way to get 16 GB of CUDA memory. The 128-bit bus means it loads big models it then runs slowly — capacity without the bandwidth to exploit it.
- VRAM
- 16 GB
- Bandwidth
- 448 GB/s
- GPU
- GB206 Blackwell, 4,608 CUDA cores
- Runs up to
- 14B at Q4, at roughly half the token rate of a 5070 Ti
Strengths
- 16 GB at the lowest tier NVIDIA offers it in
- 180 W runs happily on a modest PSU and in a small case
- Strong fit for an always-on inference box
Trade-offs
- 448 GB/s is a genuine bottleneck: roughly half a 5070 Ti
- The 8 GB variant shares the name and is a different product — check carefully
16 GB of CUDA memory at the bottom of the range, and an honest caveat.
For language models we describe this card as capacity without the bandwidth to exploit it. For image generation the bandwidth criticism largely evaporates — but a different one replaces it. At 4,608 CUDA cores it has roughly half the shader throughput of the 5070 Ti, and since diffusion is compute-bound, that is roughly half the images per minute.
So it is not the bargain here that its memory suggests. It is a perfectly capable card that will make you wait, and at 180 W it fits anywhere.
Check the variant. The 8 GB card shares the model name and is a genuinely different product.
Best used value
Most VRAM per Dollar
NVIDIA GeForce RTX 3090
3.7/5AI Gear Stack review score: 3.7 out of 5
Best for 24 GB of CUDA memory for the least money anywhere
The best capability per dollar in local AI, and it has been for years. 24 GB at 936 GB/s runs 32B-class models, and the used market prices it well below anything new with that much memory.
- VRAM
- 24 GB
- Bandwidth
- 936 GB/s
- GPU
- GA102 Ampere, 10,496 CUDA cores
- Runs up to
- 32B at Q4 with modest context
Strengths
- 24 GB reaches 32B-class models — a tier most current mid-range cards cannot touch
- 936 GB/s is faster than several newer, more expensive cards
- Memory capacity ages far more slowly than compute, and this card is the proof
- Widely available second-hand and well understood
Trade-offs
- Second-hand only, so no warranty and variable history
- No FP8 or FP4 support — newer numeric formats pass it by
- High idle power compared with current cards
- 350 W and a large physical footprint
- GDDR6X modules on this generation run hot; check thermal pad condition
24 GB and 10,496 CUDA cores, at used-market prices.
For image generation it remains genuinely capable: enough memory for Flux at reduced precision with room to spare, and enough compute to be quick. The newer cards are faster per watt and support newer numeric formats that some tooling now uses, but nothing about this card is a bottleneck for mainstream diffusion work.
The usual used-market caveats apply — no warranty, unknown history, and this generation’s GDDR6X modules ran hot, so ask about thermal pad condition.
Where the CUDA gap is widest
We say elsewhere that AMD is competitive for text generation and not for images. This is the guide where that matters, so it is worth being specific about why.
Diffusion tooling lives in ComfyUI custom nodes and in a fast-moving ecosystem of samplers, optimisation extensions and quantisation kernels. Nearly all of it is written against CUDA by people with NVIDIA hardware, and much of it reaches ROCm late or not at all.
Compounding it: image generation is compute-bound and fits in 16 GB, so AMD’s usual advantage — more memory per pound — buys you nothing here. You would be giving up ecosystem support in exchange for capacity the workload does not use.
Apple silicon works through PyTorch’s Metal backend and Core ML. It generates images, more slowly than an equivalent NVIDIA card, and with the same custom-node gaps.
For image generation specifically, our recommendation is unambiguous: buy NVIDIA. The full reasoning is in NVIDIA vs AMD for Local AI.
Training LoRAs
Worth a note because it changes the memory calculation.
Training a LoRA holds the base model, the adapter weights, gradients and optimiser state simultaneously. For SDXL that is comfortable on 16 GB with sensible settings; for the 12-billion-parameter models it is not, and 24 GB becomes the practical floor.
If LoRA training is a regular part of your work rather than an occasional experiment, weight memory more heavily than this guide otherwise suggests — the calculation moves back toward the language-model logic.
Getting more from the card you have
Before buying, four things that reclaim meaningful headroom:
- Enable tiled VAE decoding. Removes the high-resolution failure described above.
- Use a quantised checkpoint. Flux at fp8 halves the memory for a quality difference most people cannot identify in normal use.
- Check your attention implementation. Modern PyTorch includes an efficient attention path by default; older installations pinned to superseded libraries leave performance unclaimed.
- Close everything else. A browser with hardware acceleration and a desktop compositor can hold a surprising amount of VRAM.
Common questions
How much VRAM do I need for Stable Diffusion?
8 GB runs SD 1.5 and SDXL at normal resolutions. 16 GB is the sensible floor for the large modern models such as Flux at reduced precision, and for comfortable batch sizes. 24 GB is for Flux at full fp16, video generation, or training LoRAs.
Why is the RTX 5080 better than the 5070 Ti here when your other guides say the opposite?
Because the bottleneck is different. Language model generation is memory-bandwidth-bound, and those two cards have nearly identical bandwidth. Image generation is compute-bound, and the 5080 has roughly twice the shader throughput. Same cards, different workload, different answer.
Is the RTX 5090 worth it for image generation?
Only for the largest models. It is the fastest card per image by a clear margin, and for SDXL work you are paying for 32 GB the workload never touches. If you run Flux at full precision, generate video, or keep long ComfyUI chains resident, it removes every constraint at once.
Can I use an AMD card for Stable Diffusion?
It works, with more setup and a higher chance that a given custom node or extension will not. This is the workload where the CUDA gap is widest — and since diffusion is compute-bound and fits in 16 GB, AMD’s memory advantage does not offset it.
Why do I run out of memory at the end of a generation?
The VAE decode step, which converts the latent back into pixels, spikes memory well above the generation that preceded it — and at high resolution it is the operation that fails. Enable tiled VAE decoding, which every mainstream interface offers. It is slightly slower and removes the ceiling.
Does memory bandwidth matter at all for image generation?
Very little, once the model fits. That is what makes this workload different: the model is loaded once and the work is repeated arithmetic on data already in fast memory. It is why a card we describe as bandwidth-starved for language models can be perfectly good here.
Is a used RTX 3090 still good for this?
Yes. 24 GB and 10,496 CUDA cores handle mainstream diffusion work comfortably, including Flux at reduced precision. Newer cards are faster per watt and support numeric formats some tooling now uses, but nothing about it is a bottleneck. Check thermal pad condition — this generation’s memory modules ran hot.
Continue your research
- Stable Diffusion GPU Requirements — the full requirements picture, including CPU and RAM
- Best GPUs for Local AI — the same cards judged for language models, where the answer differs
- NVIDIA vs AMD for Local AI — why this is the workload where the gap is widest
- Best GPUs Under $1,000 for AI — if the budget is the constraint