NVIDIA GeForce RTX 3090 Review: Still the Answer, and Not Forever

Eleven articles on this site lean on this card. Twenty-four gigabytes for the least money is still unbeaten — and the risks around it have been growing while the recommendation stayed the same.

  • AI performance
  • General performance
  • Value
  • Build quality
  • Thermals
  • Noise
  • Power efficiency
  • Connectivity
3.7/5Overall Score

The verdict

3.7 /5

Still the cheapest route to 24 GB and to 32B-class models, five years on — and the risks around it have been growing while that recommendation stayed the same. Buy it for the capacity, with new thermal pads, a large power supply and no warranty priced in. The missing FP8 support is what will eventually retire it.

Scored against our methodology

AI performance
4
General performance
3.8
Value
4.5
Build quality
3.2
Thermals
3
Noise
3.4
Power efficiency
2.6
Connectivity
3.8

Buy it if

Running 32B-class models locally for the least money, when you can accept a used card with no warranty and have the power headroom and case clearance for it.

Skip it if

Anyone who wants a warranty or low idle power; an always-on machine, where 350 W compounds; or anyone whose tooling depends on FP8.

Strengths

  • 24 GB of CUDA memory for less than anything else, new or used
  • 936 GB/s beats every new card near its price
  • Reaches 32B-class models, which no 16 GB card can load at any setting
  • The last GeForce card with NVLink — the reason to pair these rather than 4090s
  • Only 8% slower than a 4090 for language model generation
  • Still a capable card for image generation and general work

Trade-offs

  • No FP8 or FP4 support, and tooling increasingly assumes FP8
  • GDDR6X runs hot and pads have hardened on every five-year-old unit
  • Transient spikes mean an 850 W supply for a 350 W card
  • Poor idle draw, which compounds on an always-on machine
  • Used-only, no warranty, no RMA, and supply thins rather than improves
  • Every unit has five or more years of unknown history

Eleven articles on this site lean on the RTX 3090. It leads Best GPUs Under $1,000, it is the card both the 5070 Ti and 5060 Ti reviews use as the argument against themselves, and the dual-GPU 70B machine is built around a pair of them.

That makes this the review with the most to lose. If the recommendation is wrong, a lot of the site is wrong with it.

It is not wrong. It is getting older in a way the recommendation has not been saying out loud, and this page is where that gets said.

Full specifications — NVIDIA GeForce RTX 3090
Identity
ManufacturerNVIDIA
Product familyGeForce RTX 30 Series
ModelRTX 3090
Form factorGraphics card
Release year2020
Price classMid-range ($500–$1,000)
Compute
GPUGA102 Ampere, 10,496 CUDA cores
GPU architectureAmpere (GA102)
VRAM (GB)24
VRAM typeGDDR6X, 384-bit
Memory bandwidth (GB/s)936
Compute capability noteCUDA compute 8.6; no FP8 or FP4 support
Connectivity & physical
Power draw350 W total board power; 850 W system PSU recommended
Workload suitability
Local LLMsExcellent
OllamaExcellent
LM StudioExcellent
Stable Diffusion / ComfyUIGood
Fine-tuningGood
HomelabWorkable with caveats
Developer workstationGood
Largest comfortable model32B at Q4 with modest context

Why it is still the answer

One number, and it has not been beaten for the money in five years.

RTX 3090 (used)RTX 5070 Ti (new)RTX 5060 Ti 16GB (new)
VRAM24 GB16 GB16 GB
Memory bandwidth936 GB/s896 GB/s448 GB/s
CUDA cores10,4968,9604,608
Board power350 W300 W180 W

Capacity decides which models you can run and bandwidth decides how fast. This card leads on both against every new card at anything like its price, and the capacity lead is the one that matters most, because it is a gate rather than a gradient.

24 GB reaches 32B-class models. 16 GB does not, at any setting. That is the whole recommendation in one line, and no amount of newer silicon at 16 GB changes it.

Model classAt Q4_K_MOn 24 GB
8B~5 GBFast, long context available
14B~8 GBComfortable
32B~19 GBFits, with a working context window
70B~42 GBNeeds a second card

It is also the last GeForce card with NVLink, which is why the dual-GPU architecture specifies it over the faster 4090 — a 112.5 GB/s direct link is exactly what tensor-parallel inference wants, and every card since has gone without.

What the recommendation has been under-stating

Five things, and they have all been getting worse each year while the headline advice stayed the same.

It has no FP8, and that will eventually matter

Ampere is CUDA compute capability 8.6. It has no FP8 tensor support and no FP4 — those arrived with Ada and Blackwell respectively.

Today this costs little: the GGUF quantisations most people run were designed around integer formats and work fine. But an increasing amount of tooling assumes FP8 is available, particularly in serving stacks and newer quantisation schemes, and the direction of travel is one way. This is the thing that will eventually retire the card, not its speed.

Nobody can tell you when. What you can say is that a card bought today on a five-year-old architecture has a shorter useful life ahead of it than the specifications alone suggest.

The memory runs hot, by design

The GDDR6X on this generation runs hot, and on many designs the modules sit on the back of the board where the cooler barely reaches them. What matters is not the core temperature but the memory junction temperature, which throttles around 110 °C.

On a five-year-old card the thermal pads have hardened. That is the normal condition of one, not evidence of a bad example — but it means a repaste and new pads should be in the price you are willing to pay, not a surprise afterwards.

The transients are why an 850 W supply is specified

350 W of board power with an 850 W recommendation looks like over-specification and is not. This generation draws brief current spikes well above its rated power — microsecond excursions that never appear on a wall meter and are exactly what trips a supply’s over-current protection.

A machine that reboots under load rather than crashing is almost always this. Size for the transients, and prefer an ATX 3.1 unit.

Idle draw is poor

Modern cards idle low. This one does not, and on an always-on machine that difference compounds into something you pay for continuously — which is the entire thesis of the always-on server, and the reason that architecture reaches for a 180 W card instead.

Every unit is now five years old

There is no new supply and there will not be. Every card on the market has five or more years on it, some of it spent mining, some spent running AI workloads hot for months at a time. Availability thins rather than improves, and prices on a discontinued card in demand do not behave sensibly.

Buying one, specifically

Since used is the only option, this is the operative section.

Ask what it did. Mining and sustained AI work are both harder on a card than gaming. A seller who answers the question straightforwardly is worth more than one who does not.

Ask for photographs of the PCB and the connectors, not just the shroud. Look for discolouration around the power connectors and any sign the card has been apart.

Budget for pads and paste. Roughly the price of a decent meal, and it addresses the card’s main design weakness rather than merely hoping.

Confirm your power supply has three 8-pin PCIe connectors — many of these cards need them, and adapters daisy-chained off two are how people discover their supply is inadequate.

Confirm case clearance with the front fans fitted. These are large cards and the headline figure is routinely 20–30 mm optimistic.

Accept that there is no warranty and no RMA. A failure is a return to the used market, not a replacement. That belongs in the price.

Against a 4090

The 4090 review reaches the same conclusion from the other side and it is worth restating here, because it is the comparison people expect to go the other way.

Both cards hold 24 GB. The 4090 reads it 8% faster — 1,008 GB/s against 936 — which on a 32B model is roughly 54 tokens per second against 50, a gap nobody notices in conversation. What the 4090 adds is 56% more shaders, and language model generation does not use shaders.

For chat, these are nearly the same card and one costs considerably more. For diffusion, fine-tuning and long prompts, the 4090 is far ahead. Buy on which of those you actually do.

The score, and how it is calculated

This card totals 3.7, which places it fourth of the seven GPUs reviewed here — above the RTX 5080 and both 60-class cards, and below the three that are simply more capable.

It did not always. Under a flat average of eight criteria it scored lowest of all of them, because three of those criteria — thermals, noise and power efficiency — reward a small, cool, low-power card, and this is emphatically not one. The result was that the two weakest cards for local AI both out-totalled it, which is a scale measuring the wrong thing on a site about AI hardware.

The criteria are now weighted. AI performance carries four times the weight of an ordinary criterion and value twice, while noise and connectivity carry half. The methodology page sets out the full table and the reasoning.

The effect on this card is exactly what you would expect: its 4.0 for AI performance and 4.5 for value now count for what they are worth, and its genuinely poor 2.6 for power efficiency counts for less without disappearing.

Even so, read the individual criteria rather than the total. A recommendation answers one question, and when the question is “what runs 32B models for the least money”, one row on this page decides it and the others are things you accept.

Who should buy one

Buy it if you want 32B-class models locally for the least money, you can accept a used card with no warranty, and you have the power headroom and the case for it.

Buy a 5070 Ti instead if you want a warranty, lower power and a card nobody has run hot for four years — accepting a 14B ceiling in exchange.

Buy a 5060 Ti 16GB instead if power or size is genuinely what limits you. 180 W against 350 W is the entire argument, and it wins when that constraint binds.

Buy a 5090 instead if the budget reaches. 32 GB and 1,792 GB/s is another tier, and it is new.

Buy two of these if the target is 70B-class models, which is what the dual-GPU architecture exists for — and read the section on NVLink before deciding, because it is the reason to prefer this card over a newer one for that job.

Frequently asked questions

Is the RTX 3090 still worth buying in 2026?

For local language models, yes — it remains the cheapest route to 24 GB of CUDA memory, and 24 GB is what reaches 32B-class models where 16 GB cards stop at 14B. What has changed is the risk around it rather than the capability: every unit is five or more years old, there is no warranty, the thermal pads will need replacing, and the lack of FP8 support will eventually matter. Buy it for the capacity, with those costs priced in.

RTX 3090 or RTX 5070 Ti?

The 3090 on capability and the 5070 Ti on ownership. The older card has 24 GB against 16, 936 GB/s against 896, and 17% more shaders — it reaches a tier of model the newer one cannot load. The 5070 Ti gives you a warranty, 50 W less, FP4 support and a card nobody has run hot for four years. If 32B models matter, that decides it; if they do not, the newer card is the easier life.

RTX 3090 or RTX 4090?

For chat, they are nearly the same card. Both hold 24 GB and the 4090 reads it only 8% faster — roughly 54 tokens per second against 50 on a 32B model, which nobody notices. The 4090 adds 56% more shaders, and language model generation does not use shaders. Where it pulls decisively ahead is image generation, fine-tuning and long prompts, all of which are compute-bound. Buy on which of those you actually do.

Does the missing FP8 support matter?

Little today and more each year. The GGUF quantisations most people run were designed around integer formats and work perfectly well on Ampere. But serving stacks and newer quantisation schemes increasingly assume FP8 is available, and the direction of travel is one way. This is what will eventually retire the card — not its speed, which is still competitive. Nobody can say when, only that a five-year-old architecture has less life ahead of it than its specifications suggest.

What should I check when buying one used?

Ask what the card did — mining and sustained AI work are both harder on it than gaming. Ask for photographs of the PCB and the power connectors rather than just the shroud, and look for discolouration. Confirm your supply has three 8-pin PCIe connectors, since many of these cards need them. Confirm case clearance with the front fans fitted. And budget for new thermal pads and paste, because after five years they have hardened and this generation’s memory runs hot.

Why does a 350 W card need an 850 W power supply?

Because you are sizing for transients rather than average draw. This generation pulls brief current spikes well above its rated board power — microsecond excursions that never show on a wall meter and are exactly what trips a supply’s over-current protection. A machine that reboots under load rather than crashing is almost always an undersized or older supply, and an ATX 3.1 unit is specified to tolerate that behaviour.

How is the overall score calculated?

It is a weighted mean rather than a flat average. AI performance carries four times the weight of an ordinary criterion and value twice, while noise and connectivity carry half — the full table is on the methodology page. A flat average was measuring the wrong thing for a site about AI hardware, because thermals, noise and power efficiency between them outweighed the one criterion most readers come here for. Even weighted, the individual rows carry more meaning than the total.

Did AI Gear Stack test this card?

No. This assessment is drawn from manufacturer specifications and published architectural detail, as stated at the top of the page. Throughput figures are calculated from the memory bandwidth specification rather than measured. Acoustics, sustained thermal behaviour and the condition of any particular second-hand card are the questions this method cannot answer, and we say so rather than guessing.

As an Amazon Associate, AI Gear Stack earns from qualifying purchases. Amazon and the Amazon logo are trademarks of Amazon.com, Inc. or its affiliates.