This tier is unusual because the best answer is frequently not a current product.
Memory capacity is what determines which models you can run, and capacity has aged far more slowly than compute. A flagship from two generations ago has more of it than most new cards at this price — which makes the used market the central fact of shopping here, not a footnote.
Our picks at a glance
-
Best Overall
NVIDIA GeForce RTX 3090
24 GB at 936 GB/s — the capability step no new card here reaches
-
Best Premium
NVIDIA GeForce RTX 5070 Ti
The best new card at this budget, with a warranty
-
Most VRAM per Dollar
AMD Radeon RX 7900 XTX
24 GB new, if your workload is text generation on Linux
-
Best Low-Power Option
NVIDIA GeForce RTX 5060 Ti 16GB
16 GB at 180 W for an always-on box
-
Best Budget
NVIDIA GeForce RTX 4060 Ti 16GB
The cheapest 16 GB — and the slowest worth considering
A note on price
You will not find dollar figures against products anywhere on this site — our editorial policy commits us to never publishing a price we cannot verify at the moment you read it, and a figure captured at writing is wrong within days.
So this guide works in capability tiers. The budget in the title is yours; check current prices yourself against the named cards, because at this tier the ordering genuinely changes with the market.
What this budget reaches
| Capability | Available under $1,000? |
|---|---|
| 16 GB, new | Comfortably — several options |
| 24 GB, used | Yes, and this is the interesting one |
| 24 GB, new | No |
| 32 GB, any | No |
| 14B models at good speed | Yes |
| 32B models at 4-bit | Only with 24 GB, so only used |
| Image generation, any mainstream model | Yes |
| 70B-class models | No |
The row that shapes everything is the third against the second. New cards at this price top out at 16 GB. Used cards reach 24. And 24 GB is the difference between 14B-class models and 32B-class ones, which is a real capability step rather than a marginal one.
Bus width matters more here than anywhere
At the top of the market every card has a wide memory bus. At this tier they do not, and two cards with the same 16 GB can be a factor of three apart on speed.
| Card | VRAM | Bus | Bandwidth |
|---|---|---|---|
| RTX 3090 (used) | 24 GB | 384-bit | 936 GB/s |
| RX 7900 XTX | 24 GB | 384-bit | 960 GB/s |
| RTX 5070 Ti | 16 GB | 256-bit | 896 GB/s |
| RTX 5060 Ti 16GB | 16 GB | 128-bit | 448 GB/s |
| RTX 4060 Ti 16GB | 16 GB | 128-bit | 288 GB/s |
Since token generation is bandwidth-bound, that last column is roughly proportional to how fast a model runs once it fits. The 4060 Ti loads exactly what the 5070 Ti loads and runs it at about a third of the speed.
Read the bus width, not just the capacity. It is the specification that separates a bargain from a disappointment at this tier.
The specifications
| Product | VRAM | Bandwidth | Power | Runs up to | Local LLMs | Where to buy |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 3090 NVIDIA | 24 GB | 936 GB/s | 350 W total board power; 850 W system PSU recommended | 32B at Q4 with modest context | Excellent | Check Price on Amazon NVIDIA GeForce RTX 3090 at Amazon — opens in a new tab |
| AMD Radeon RX 7900 XTX AMD | 24 GB | 960 GB/s | 355 W total board power; 800 W system PSU recommended | 32B at Q4 under llama.cpp or Ollama | Good | Check Price on Amazon AMD Radeon RX 7900 XTX at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5070 Ti NVIDIA | 16 GB | 896 GB/s | 300 W total board power; 750 W system PSU recommended | 14B at Q4 with comfortable context headroom | Good | Check Price on Amazon NVIDIA GeForce RTX 5070 Ti at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 5060 Ti 16GB NVIDIA | 16 GB | 448 GB/s | 180 W total board power; 600 W system PSU recommended | 14B at Q4, at roughly half the token rate of a 5070 Ti | Workable with caveats | Check Price on Amazon NVIDIA GeForce RTX 5060 Ti 16GB at Amazon — opens in a new tab |
| NVIDIA GeForce RTX 4060 Ti 16GB NVIDIA | 16 GB | 288 GB/s | 165 W total board power; 550 W system PSU recommended | 14B at Q4, slowly | Workable with caveats | Check Price on Amazon NVIDIA GeForce RTX 4060 Ti 16GB at Amazon — opens in a new tab |
The recommendations
Best overall
Best Overall
NVIDIA GeForce RTX 3090
3.7/5AI Gear Stack review score: 3.7 out of 5
Best for 24 GB of CUDA memory for the least money anywhere
The best capability per dollar in local AI, and it has been for years. 24 GB at 936 GB/s runs 32B-class models, and the used market prices it well below anything new with that much memory.
- VRAM
- 24 GB
- Bandwidth
- 936 GB/s
- GPU
- GA102 Ampere, 10,496 CUDA cores
- Runs up to
- 32B at Q4 with modest context
Strengths
- 24 GB reaches 32B-class models — a tier most current mid-range cards cannot touch
- 936 GB/s is faster than several newer, more expensive cards
- Memory capacity ages far more slowly than compute, and this card is the proof
- Widely available second-hand and well understood
Trade-offs
- Second-hand only, so no warranty and variable history
- No FP8 or FP4 support — newer numeric formats pass it by
- High idle power compared with current cards
- 350 W and a large physical footprint
- GDDR6X modules on this generation run hot; check thermal pad condition
The best capability per dollar in local AI, and it has held that position for years.
24 GB at 936 GB/s runs 32B-class models at 4-bit — a tier that no new card at this budget reaches — and it does so quickly, because the 384-bit bus gives it more bandwidth than several newer and more expensive cards.
This is the clearest illustration of the principle this site keeps returning to: memory capacity ages slowly, compute ages quickly. A card from 2020 outperforms current mid-range hardware at the thing you actually want to do, because model sizes have only grown and 24 GB is still 24 GB.
What you accept in return: no warranty, an unknown history, no FP8 or FP4 support, high idle power, and 350 W in a physically large card. This generation’s GDDR6X modules also ran hot, so ask about thermal pad condition and be wary of a card that has spent its life at full load.
Best new card
Best Premium
NVIDIA GeForce RTX 5070 Ti
3.9/5AI Gear Stack review score: 3.9 out of 5
Best for The cheapest sensible entry to 16 GB of fast NVIDIA VRAM
Within about 7% of a 5080 on memory bandwidth with the same 16 GB, at meaningfully lower power and price. For most people choosing between the two, this is the better buy.
- VRAM
- 16 GB
- Bandwidth
- 896 GB/s
- GPU
- GB203 Blackwell, 8,960 CUDA cores
- Runs up to
- 14B at Q4 with comfortable context headroom
Strengths
- Same 16 GB and near-identical bandwidth to the RTX 5080
- 300 W fits comfortably in an existing mid-range build
- Strong price-to-capability ratio for image generation
Trade-offs
- 16 GB caps language-model ambitions at roughly 14B
- No capability advantage over the 5080 — only a price one
If buying used is not for you, this is the top of what this budget reaches new, and it is a genuinely good card.
16 GB at 896 GB/s covers 14B-class models with long context and runs image generation quickly. At 300 W it drops into an existing build without a power supply upgrade, which at this budget is a real saving rather than a footnote.
Its limitation is the one it shares with every 16 GB card: 32B models are out of reach. That is the gap the used 3090 fills.
Best value on capacity
Most VRAM per Dollar
AMD Radeon RX 7900 XTX
3.9/5AI Gear Stack review score: 3.9 out of 5
Best for 24 GB of VRAM for text generation, at a lower price than the NVIDIA equivalent
24 GB and 960 GB/s for less money than a comparable NVIDIA card. The hardware is not the problem — the software ecosystem is, and how much that costs you depends entirely on your workload.
- VRAM
- 24 GB
- Bandwidth
- 960 GB/s
- GPU
- Navi 31 RDNA 3, 96 compute units
- Runs up to
- 32B at Q4 under llama.cpp or Ollama
Strengths
- 24 GB at a considerably lower price than NVIDIA charges
- 960 GB/s matches an RTX 5080 despite the older architecture
- Text generation through llama.cpp and Ollama is well supported and fast
Trade-offs
- ROCm support is narrower and more fragile than CUDA
- Image generation and fine-tuning workflows frequently assume CUDA
- Expect to spend time on setup that an NVIDIA card would not require
24 GB and 960 GB/s, new, typically below this budget — the only way to get that much memory without buying second-hand.
For text generation through Ollama or llama.cpp this is a strong proposition, and the hardware genuinely competes. For image generation, fine-tuning, or anything that assumes CUDA, expect friction that the memory advantage does not offset. Our NVIDIA vs AMD comparison works through exactly where that line falls.
If you run language models on Linux and enjoy getting things working, this is excellent value. If you want to install a runtime and get on with your work, it is not.
Best low-power option
Best Low-Power Option
NVIDIA GeForce RTX 5060 Ti 16GB
3.6/5AI Gear Stack review score: 3.6 out of 5
Best for Getting 16 GB of CUDA VRAM into a small or low-power machine
The cheapest way to get 16 GB of CUDA memory. The 128-bit bus means it loads big models it then runs slowly — capacity without the bandwidth to exploit it.
- VRAM
- 16 GB
- Bandwidth
- 448 GB/s
- GPU
- GB206 Blackwell, 4,608 CUDA cores
- Runs up to
- 14B at Q4, at roughly half the token rate of a 5070 Ti
Strengths
- 16 GB at the lowest tier NVIDIA offers it in
- 180 W runs happily on a modest PSU and in a small case
- Strong fit for an always-on inference box
Trade-offs
- 448 GB/s is a genuine bottleneck: roughly half a 5070 Ti
- The 8 GB variant shares the name and is a different product — check carefully
16 GB of CUDA memory at the bottom of the range, at 180 W.
At 448 GB/s it is half the speed of the 5070 Ti on language models. For an always-on inference box, a homelab node, or a small case where power and heat are the binding constraints, that trade is sensible. For a machine you sit in front of, it is not.
The cheapest 16 GB
Best Budget
NVIDIA GeForce RTX 4060 Ti 16GB
3.2/5AI Gear Stack review score: 3.2 out of 5
Best for Getting 16 GB into a small, low-power or second-hand build cheaply
Sixteen gigabytes at 288 GB/s. It will load what a 5070 Ti loads and run it at roughly a third of the speed — capacity without the bandwidth to use it, which is a defensible trade only at the right price.
- VRAM
- 16 GB
- Bandwidth
- 288 GB/s
- GPU
- AD106 Ada Lovelace, 4,352 CUDA cores
- Runs up to
- 14B at Q4, slowly
Strengths
- 16 GB at the bottom of the market
- 165 W runs on almost any power supply, in almost any case
- A capable image-generation card, where bandwidth matters far less
- Widely available new and used
Trade-offs
- 288 GB/s is the lowest of any 16 GB card worth considering
- PCIe ×8 halves the offload path when a model does not fit
- The 8 GB variant shares the name — check carefully
- Superseded by the RTX 5060 Ti 16GB, which is faster for similar money
Included for completeness and with a recommendation attached: buy the RTX 5060 Ti 16GB instead if you can.
At 288 GB/s this is the slowest 16 GB card worth considering, and it uses only eight PCIe lanes — which halves the offload path on the occasions a model does not fit. The newer 5060 Ti is meaningfully faster for similar money.
It remains a reasonable image-generation card, where bandwidth matters far less, and a reasonable purchase if the price gap is large.
What to avoid
8 GB cards. They run 7B models and nothing beyond, and 8 GB was already tight before the current generation of models. The money saved buys you a card you will replace.
The name collision. Both the 4060 Ti and 5060 Ti ship in 8 GB and 16 GB variants under the same model name. Check the listing carefully — this is the single most common ordering mistake at this tier.
Anything without a stated bus width. If a listing does not say, it is usually 128-bit.
Buying used, carefully
Since the best answer here is frequently second-hand, the checklist is worth having:
- Inspect the power connector for discolouration or melting. On 12VHPWR-equipped cards this was a genuine failure mode.
- Ask what it was used for. A card that spent three years mining or running sustained compute has had a harder life than a gaming card, and thermal pads and fan bearings both age.
- Buy from a seller who accepts returns. A platform with buyer protection is worth the small premium over a private sale.
- Test it within the return window. Run a sustained load for an hour and watch memory temperatures, not just core temperatures — on this generation the memory is what fails first.
- Budget for the power supply. A 350 W card in a machine built around a 550 W unit means buying a supply too.
Common questions
Is a used RTX 3090 better than a new mid-range card?
For local AI, usually yes. 24 GB reaches 32B-class models that no 16 GB card can load, and its 936 GB/s is faster than several newer cards. You give up the warranty, FP8 and FP4 support, and low idle power. It is the clearest example of memory capacity ageing better than compute.
What is the best new GPU under $1,000 for AI?
The RTX 5070 Ti. 16 GB at 896 GB/s covers 14B-class models with long context, runs image generation quickly, and at 300 W fits an existing build without a power supply upgrade.
Why are two 16 GB cards so different in speed?
Memory bus width. A 256-bit card carries roughly twice the bandwidth of a 128-bit one at similar memory speeds, and token generation is bandwidth-bound. The RTX 4060 Ti 16GB loads exactly what a 5070 Ti loads and runs it at about a third of the pace.
Should I buy an 8 GB card to save money?
No. 8 GB runs 7B models and nothing beyond, and that ceiling was already tight before the current generation of models. The saving buys a card you will replace rather than one you will keep.
Is the RX 7900 XTX good value here?
On hardware, genuinely — 24 GB and 960 GB/s, new, at this budget. On software it depends entirely on your workload: strong for text generation through Ollama or llama.cpp, and meaningfully worse for image generation, fine-tuning or anything assuming CUDA.
What should I check when buying a used GPU?
The power connector for heat damage, what the card was used for, and that the seller accepts returns. Then test it under sustained load within the return window and watch memory temperatures rather than core temperatures — on the RTX 30 series the GDDR6X modules are what fail first.
Continue your research
- Best GPUs for Local AI — the full range, unconstrained by budget
- How Much VRAM Do You Need for Local AI? — establishing your actual requirement
- Best AI Workstations Under $2,000 — the machine around the card
- Best GPUs for Stable Diffusion — if images rather than text are the workload