NVIDIA GeForce RTX 4090 Review: Brilliant at the Work Language Models Do Not Do

Twenty-four gigabytes and a 56% shader advantage over a used 3090 — almost none of which language model generation uses. An excellent card whose case rests on what else you run.

  • AI performance
  • General performance
  • Value
  • Build quality
  • Thermals
  • Noise
  • Power efficiency
  • Connectivity
4/5Overall Score

The verdict

4 /5

A genuinely excellent card that is hard to recommend for language models specifically, because almost everything it does better than a used 3090 is compute — and generation barely uses compute. Buy it for diffusion, fine-tuning and prefill; buy a 3090 for chat.

Scored against our methodology

AI performance
4.2
General performance
4.5
Value
2.9
Build quality
4
Thermals
4.4
Noise
4.3
Power efficiency
4.1
Connectivity
3.8

Buy it if

Compute-bound work — image generation, fine-tuning, long-context prompt processing — where you want 24 GB without paying 5090 money or building a machine around 575 W.

Skip it if

Running language models and little else, where a used 3090 gives you the same 24 GB and about 92% of the generation speed for considerably less; or anyone who needs 70B-class models, which 24 GB cannot reach.

Strengths

  • 24 GB reaches 32B-class models, which every 16 GB card stops short of
  • 56% more shaders than a 3090 — decisive for image generation and fine-tuning
  • Markedly faster prompt processing on long documents
  • FP8 tensor support, which Ampere lacks
  • 450 W is far easier to build around than the 5090's 575 W
  • Large coolers and good sustained thermal behaviour

Trade-offs

  • Only 8% faster than a used 3090 for language model generation, at a much higher price
  • Out of production: no warranty, no RMA, used market only
  • 12VHPWR connector with a documented seating-failure history to inspect
  • 100 W more than a 3090 for that 8%
  • The RTX 5090 offers 78% more bandwidth and eight more gigabytes
  • DisplayPort 1.4a rather than 2.1b

The RTX 4090 is a genuinely excellent graphics card that is difficult to recommend for language models, and the reason is specific enough to be worth stating in one sentence: almost everything it does better than a used RTX 3090 is compute, and language model generation barely uses compute.

That is not a criticism of the card. It is a description of a workload, and it means the 4090’s case rests almost entirely on what else you intend to do with it.

Full specifications — NVIDIA GeForce RTX 4090
Identity
ManufacturerNVIDIA
Product familyGeForce RTX 40 Series
ModelRTX 4090
Form factorGraphics card
Release year2022
Price classFlagship ($3,500+)
Compute
GPUAD102 Ada Lovelace, 16,384 CUDA cores
GPU architectureAda Lovelace (AD102)
VRAM (GB)24
VRAM typeGDDR6X, 384-bit
Memory bandwidth (GB/s)1008
Connectivity & physical
Power draw450 W total board power; 850 W system PSU recommended
Workload suitability
Local LLMsExcellent
OllamaExcellent
Stable Diffusion / ComfyUIExcellent
Fine-tuningGood
HomelabWorkable with caveats
Developer workstationExcellent
Largest comfortable model32B at Q4 with modest context, or 14B at Q8

The comparison that decides it

The 4090 has been out of production for some time, so buying one means the used market — which puts it directly against the other 24 GB card on that market.

RTX 3090RTX 4090Difference
VRAM24 GB GDDR6X24 GB GDDR6XNone
Memory bandwidth936 GB/s1,008 GB/s+8%
CUDA cores10,49616,384+56%
Board power350 W450 W+100 W
ArchitectureAmpereAda LovelaceTwo generations

Read the first two rows together, because they are what a language model actually cares about. Generation is memory-bandwidth-bound: the card reads its weights out of memory to produce each token, capacity decides which models fit, and bandwidth decides how quickly they run.

Both cards hold 24 GB. The 4090 reads it 8% faster. On a 32B model at Q4 that is the difference between roughly 50 and roughly 54 tokens per second — a gap you would struggle to notice in conversation, for a card that costs considerably more.

The +56% shader count is real and it is large. It simply does not appear in that number.

Where the extra silicon does show up

Three workloads, and if any of them is yours the card’s case changes completely.

Image generation. Diffusion is compute-bound, not bandwidth-bound — the Stable Diffusion guide works through why. Here the 56% shader advantage is close to the whole story, and the 4090 is far ahead of a 3090 rather than marginally. It also has FP8 tensor support, which Ampere lacks and which some tooling now uses.

Prompt processing. Prefill — reading everything you pasted before the first token appears — is compute-bound in a way that decode is not. Feed a long document to both cards and the 4090 starts answering appreciably sooner, even though it then streams at almost the same speed.

Fine-tuning. Forward and backward passes over batches are dense arithmetic. The fine-tuning workstation explains why compute stops being irrelevant the moment you are changing weights rather than reading them, and this is the card’s strongest showing against its predecessor.

If you do none of those — if the machine exists to run a chat model locally — you are paying a large premium for capability that will sit idle.

What 24 GB runs

The capacity is the reason to want this card, and it is the same reason to want a 3090.

Model classAt Q4_K_MOn 24 GB
8B~5 GBVery fast, long context available
14B~8 GBComfortable, the everyday tier
32B~19 GBFits, with room for a working context window
70B~42 GBDoes not fit

That 32B row is the whole argument for 24 GB over 16. A 5080 or 5070 Ti stops at the 14B tier; this card reaches the one above it, and the jump in output quality between them is larger than any difference in tokens per second.

The 70B row is the ceiling, and no setting moves it. Reaching that tier means two cards, a 96 GB professional card, or unified memory.

Against the RTX 5090

The three 24 GB-and-above consumer cards this review sits between.
Product VRAM Bandwidth Power Runs up to Where to buy
NVIDIA GeForce RTX 3090 NVIDIA 24 GB 936 GB/s 350 W total board power; 850 W system PSU recommended 32B at Q4 with modest context Check Price on Amazon NVIDIA GeForce RTX 3090 at Amazon — opens in a new tab
NVIDIA GeForce RTX 4090 NVIDIA 24 GB 1008 GB/s 450 W total board power; 850 W system PSU recommended 32B at Q4 with modest context, or 14B at Q8 Check Price on Amazon NVIDIA GeForce RTX 4090 at Amazon — opens in a new tab
NVIDIA GeForce RTX 5090 NVIDIA 32 GB 1792 GB/s 575 W total board power; 1,000 W system PSU recommended 32B at Q4 entirely in VRAM, with room for long context Check Price on Amazon NVIDIA GeForce RTX 5090 at Amazon — opens in a new tab

The 5090 is 78% more memory bandwidth and eight more gigabytes, which are precisely the two specifications that govern language models. It is not a close contest on capability: 32 GB reaches 32B models with far more context headroom, and generation is meaningfully faster rather than marginally.

What the 4090 keeps is a 450 W board power against 575 W, which is the difference between a machine most people can build without thinking about it and one that needs the power and cooling planning the $2,000 architecture describes. It also asks less of the case and the room.

If budget were not a consideration the 5090 is the better card for this work, plainly. The 4090 exists in this conversation because used pricing sometimes says otherwise.

Buying one used

It is the only way to buy one now, so the practical caveats matter more than usual.

The 12VHPWR connector. This generation’s connector has a documented history of failures where it was not fully seated, with consequences beyond an unstable machine. Inspect the connector and the socket for discolouration or melted plastic before buying, seat it firmly until it clicks, and route the cable so it is not bent sharply immediately behind the plug. A native cable from an ATX 3.1 supply is preferable to an adapter.

Cards that have been apart. The 4090’s die has been in demand for conversion into other products, and blower-style rebuilds exist. A card that has been disassembled is not necessarily bad, but it is not the card the manufacturer shipped and it has no warranty path. Ask for photographs of the PCB and the shroud, and be wary of anything with a cooler that does not match the model name.

Thermal pads and paste. A card that has run AI workloads has run hot for long periods, which is harder on it than gaming. Budget for a repaste, and treat the seller’s willingness to answer questions about its history as part of the price.

No warranty, and no replacement path. Production has ended, so a failure is not an RMA — it is going back to the used market. That belongs in the price you are willing to pay.

Power and the rest of the machine

450 W board power, an 850 W supply recommended, and — as with every card of this era — size the supply for transient excursions rather than steady draw. Brief spikes well above rated power are what trip over-current protection on supplies that look adequate on paper.

Physically it is a large card. Three slots is common among third-party designs and lengths past 330 mm are normal, so confirm case clearance with the front fans fitted rather than against the manufacturer’s headline figure.

Displays are the one place the age shows in a small way: DisplayPort 1.4a rather than the 2.1b of the current generation. For most monitors this changes nothing; at very high resolution and refresh together it can.

Who should buy one

Buy it if your work is compute-bound — image generation, fine-tuning, long-context prompt processing — and you want 24 GB without paying 5090 money or building for 575 W. In that case it is a strong card at a sensible price, and the shader advantage over a 3090 is money well spent.

Buy a used 3090 instead if you are running language models and little else. You get the same 24 GB and about 92% of the generation speed, for considerably less. This is the recommendation most readers of this site will land on, and it is the one the arithmetic supports.

Buy a 5090 instead if language models are the point and the budget reaches. Eight more gigabytes and 78% more bandwidth are the two things that matter here, and no amount of shader count on the 4090 closes that.

Frequently asked questions

Is the RTX 4090 still worth buying for AI in 2026?

For compute-bound work, yes — image generation, fine-tuning and long-context prompt processing all use the shader advantage that makes it 56% richer than a 3090. For running language models it is harder to justify: a used 3090 holds the same 24 GB and generates at about 92% of the speed for considerably less money, because decode is bandwidth-bound and the two cards are only 8% apart on bandwidth.

RTX 4090 or RTX 3090 for local LLMs?

The 3090, for most people. Both hold 24 GB, so both reach the same models — and generation speed tracks memory bandwidth, where 1,008 GB/s against 936 GB/s is a difference you will not notice in conversation. The 4090 is the better card in every other respect and considerably more efficient per unit of work, but you are paying for compute that language model generation does not use.

RTX 4090 or RTX 5090?

The 5090, if the budget reaches and language models are the point. Eight more gigabytes and 78% more memory bandwidth are precisely the two specifications that govern this workload. The 4090 keeps a real advantage in being a 450 W card rather than a 575 W one, which asks much less of your power supply, your case and the room it sits in.

How many tokens per second will it produce?

For a 32B model at Q4 the arithmetic ceiling is roughly 50 tokens per second — 1,008 GB/s divided by about 19 GB of weights — and real throughput lands below that, because attention, sampling and the KV cache all consume bandwidth the formula ignores. A 14B model is comfortably faster than reading speed. These are calculated figures, not measurements taken here.

Can it run a 70B model?

Not usefully. A 70B model at Q4 needs roughly 42 GB before any context, so it will not fit in 24 GB. It can be made to run by offloading layers to system memory, at speeds most people abandon within a week. Reaching that tier means two cards, a 96 GB professional card, or a unified-memory machine.

What should I check when buying one second-hand?

The 12VHPWR connector and its socket, for discolouration or melted plastic — this generation has a documented history of failures where the plug was not fully seated. Whether the card has been apart, since the die has been in demand for conversion into other products and blower-style rebuilds exist. And budget for new thermal pads and paste, because a card that has run AI workloads has run hot for long periods.

What power supply does it need?

An 850 W unit is the usual recommendation for 450 W of board power, and the ATX 3.1 standard matters as much as the wattage — brief transient excursions well above rated draw are what trip protection on supplies that look adequate on paper. Prefer a native 12V-2×6 cable to an adapter, and seat it fully.

Did AI Gear Stack test this card?

No. This assessment is drawn from manufacturer specifications and published architectural detail, as stated at the top of the page. Throughput figures are calculated from the memory bandwidth specification rather than measured. Acoustics, sustained thermal behaviour and the condition of any particular second-hand card are the questions this method cannot answer, and we say so rather than guessing.

As an Amazon Associate, AI Gear Stack earns from qualifying purchases. Amazon and the Amazon logo are trademarks of Amazon.com, Inc. or its affiliates.