This is a reference architecture, not a tested build. Every component here is specified from published manufacturer documentation — socket, TDP, connector type, memory support, dimensions — and the parts are specified to work together on paper. Nobody at AI Gear Stack has assembled this machine.
That distinction matters enough to lead with, because the failures that ruin a first build are almost never visible on a spec sheet. A cooler that overhangs the first memory slot. A card the case manufacturer lists as supported that turns out to need four millimetres you do not have once the front fans are fitted. A motherboard revision shipping firmware older than the CPU it is rated for.
What is establishable from documentation is the reasoning: which tier of card to buy, why capacity beats speed for this workload, how much power headroom a transient-spiking GPU needs, and what the finished machine will actually run. That is what follows, and it ends with a checklist of what to confirm against your own basket before you order.
The two paths at this tier
-
Most VRAM per Dollar
NVIDIA GeForce RTX 3090
A used 24 GB card reaches 32B-class models, which is the biggest single jump in output quality available at this budget. No warranty, higher idle draw, and worth budgeting for new thermal pads.
-
Best Value
NVIDIA GeForce RTX 5070 Ti
Nearly identical memory bandwidth, so similar speed on models that fit in 16 GB — and clearly the better card if the machine also does image generation, which is compute-bound rather than bandwidth-bound.
A note on price
You will not find dollar figures against products anywhere on this site — our editorial policy commits us to never publishing a price we cannot verify at the moment you read it, and a figure captured at writing is wrong within days.
The number in the title is a capability tier, not a receipt. It describes a machine built around a 16–24 GB graphics card with the rest of the system specified not to bottleneck it, which is what roughly two thousand dollars has bought for the past several product cycles. Check current prices yourself. At this tier the used market moves enough that the ordering genuinely changes.
The one decision that shapes everything else
Every other choice in this build follows from how much VRAM you buy, so make that decision first.
Local language model generation is memory-bandwidth-bound: the machine reads the model’s active weights out of memory to produce each token, so capacity decides which models you can run at all, and bandwidth decides how fast. Compute barely registers. This is why an AI build looks so unlike a gaming build — you spend on the card’s memory and starve everything else, and the result is faster than a balanced machine costing the same.
That gives two credible paths at this tier.
| Path A — 24 GB, used | Path B — 16 GB, new | |
|---|---|---|
| Graphics card | RTX 3090 (used) | RTX 5070 Ti |
| VRAM | 24 GB GDDR6X | 16 GB GDDR7 |
| Memory bandwidth | ~936 GB/s | ~896 GB/s |
| Board power | 350 W | 300 W |
| Comfortably runs | 32B-class at Q4 | 14B-class at Q4 |
| Warranty | None, typically | Full |
| Idle power and noise | Higher | Lower |
The bandwidth figures are close enough that tokens per second on a model that fits in both cards will be similar. The 3090 is not the faster card in any general sense — it is two generations older and materially less efficient. What it has is eight more gigabytes, and eight gigabytes is the difference between running 14B-class models and 32B-class models, which is a larger jump in output quality than anything else available at this budget.
Take Path A if the capability step matters more to you than the warranty, and you are comfortable buying used hardware. Take Path B if it does not, or if the machine is doing image generation as much as text — diffusion is compute-bound, which inverts the comparison entirely and puts the newer card well ahead.
The rest of this specification is identical for both, except where noted.
Most VRAM per Dollar
NVIDIA GeForce RTX 3090
3.7/5AI Gear Stack review score: 3.7 out of 5
Best for 24 GB of CUDA memory for the least money anywhere
The capability choice. Twenty-four gigabytes at 936 GB/s for used-market money, at the cost of a warranty and roughly 100 W of idle-to-load inefficiency against a current card.
- VRAM
- 24 GB
- Bandwidth
- 936 GB/s
- GPU
- GA102 Ampere, 10,496 CUDA cores
- Runs up to
- 32B at Q4 with modest context
Strengths
- 24 GB reaches 32B-class models — a tier most current mid-range cards cannot touch
- 936 GB/s is faster than several newer, more expensive cards
- Memory capacity ages far more slowly than compute, and this card is the proof
- Widely available second-hand and well understood
Trade-offs
- Second-hand only, so no warranty and variable history
- No FP8 or FP4 support — newer numeric formats pass it by
- High idle power compared with current cards
- 350 W and a large physical footprint
- GDDR6X modules on this generation run hot; check thermal pad condition
NVIDIA GeForce RTX 5070 Ti
3.9/5AI Gear Stack review score: 3.9 out of 5
Best for The cheapest sensible entry to 16 GB of fast NVIDIA VRAM
The safe choice, and the right one if images matter as much as text. Same bandwidth class, eight fewer gigabytes, full warranty and materially better efficiency.
- VRAM
- 16 GB
- Bandwidth
- 896 GB/s
- GPU
- GB203 Blackwell, 8,960 CUDA cores
- Runs up to
- 14B at Q4 with comfortable context headroom
Strengths
- Same 16 GB and near-identical bandwidth to the RTX 5080
- 300 W fits comfortably in an existing mid-range build
- Strong price-to-capability ratio for image generation
Trade-offs
- 16 GB caps language-model ambitions at roughly 14B
- No capability advantage over the 5080 — only a price one
The component specification
| Component | Specification | Why |
|---|---|---|
| Graphics card | RTX 3090 24 GB (used) or RTX 5070 Ti 16 GB | The entire point of the machine |
| CPU | AMD Ryzen 7 9700X (AM5, 8 cores, 65 W) | Fast enough to never be the limit; cheap and cool |
| Motherboard | B650 ATX, 2× M.2, PCIe 5.0 ×16 slot | Full-size board for clearance and airflow headroom |
| Memory | 2 × 32 GB DDR5-6000 CL30 | Two DIMMs, not four — see below |
| Primary storage | 2 TB NVMe, PCIe 4.0 or better | Model weights are enormous and read constantly |
| Power supply | 850 W ATX 3.1, native 12V-2×6 | Transient headroom, not average draw |
| CPU cooler | 240 mm AIO or a large dual-tower air cooler | Height and memory clearance both matter |
| Case | Mid-tower rated for ≥ 360 mm GPU length | The single most commonly wrong number in a build |
Read that as a specification rather than a shopping list. The requirements — socket, capacity, connector type, clearance — are the part that holds; a named model is one example that satisfies them.
Why each part is what it is
The graphics card
Covered above, and it is roughly two-thirds of the budget. That ratio is correct for this workload and would be indefensible for almost any other.
The CPU: deliberately unremarkable
Once the model fits in VRAM, the CPU does almost nothing during generation. It feeds the GPU and gets out of the way. A Ryzen 7 9700X is specified here because it is fast in single-thread terms, runs at 65 W so it never fights the graphics card for thermal or power budget, and costs a fraction of what the equivalent enthusiast part does.
Spending more on the CPU is the most common way to build a worse AI machine at the same price. The exception is if you intend to run models larger than your VRAM, where layers spill into system memory and the CPU does have to do real work — but that path is so much slower that the right answer is almost always to run a smaller model rather than build for it.
Memory: two sticks, not four
64 GB in two DIMMs, and this is a genuine compatibility trap rather than a preference.
AM5 runs its memory controller comfortably at DDR5-6000 with two single-rank or dual-rank modules. Populate all four slots and the controller frequently will not hold that speed — DDR5-3600 to 4400 is a common landing point, sometimes needing manual tuning to boot at all. Buying 2 × 32 GB rather than 4 × 16 GB costs slightly more per gigabyte and leaves both the speed and a future upgrade path intact.
64 GB is specified rather than 32 because model files are large, several tools memory-map them, and the machine is more pleasant to use when the page cache is not constantly evicting a 20 GB file you are about to load again.
Storage: capacity first
A single 2 TB NVMe drive. Model weights are the largest files most people will ever store routinely — a handful of quantised 32B models and a couple of image models will pass a terabyte without much effort.
PCIe 4.0 is sufficient. Load time scales with sequential read, and the difference between a good 4.0 drive and a good 5.0 drive is a few seconds on a load that happens once per session. Sustained write performance and drive endurance matter more here than the headline number, which is why the guide to NVMe SSDs for AI workloads spends most of its length on the SLC cache cliff rather than on peak figures.
The board specifies two M.2 slots so the second drive is a later decision rather than a rebuild.
The power supply: sized for spikes, not averages
850 W, which looks generous against a 350 W card and a 65 W CPU, and is not.
The reason is transient behaviour. GDDR6X-era NVIDIA cards, the 3090 emphatically included, draw brief current spikes well above their rated board power — microsecond-scale excursions that never show on a wall meter but are exactly what trips a power supply’s over-current protection. A machine that reboots under load rather than crashing is nearly always this. Modern ATX 3.1 units are specified to tolerate those excursions; older units of the same wattage frequently are not, which is why the standard and not just the number is part of the specification.
Specify native 12V-2×6 on the supply side. The adapters work, but a native cable is one fewer connection to seat imperfectly, and imperfect seating on this connector is a well-documented failure mode with consequences beyond an unstable machine.
Path B’s 300 W card is comfortable on a 750 W unit of the same standard.
Cooling and case: where the specification meets reality
The CPU is 65 W and needs almost nothing. The case has to handle a card dumping 300–350 W into it for hours at a time, which is a different problem from gaming — sustained rather than bursty, and with no pauses for the thermal mass to recover.
Specify front intake fans, and confirm the case has them fitted rather than merely supported. The pattern that causes long-run throttling is a case with excellent reviews and one exhaust fan in the box.
For the cooler, a 240 mm AIO or a large dual-tower air cooler both work. Air is cheaper and has no pump to fail; check the height against the case and the memory clearance against your DIMMs, because tall heat spreaders and wide coolers are the classic collision.
What the finished machine runs
Assuming Path A’s 24 GB, and Q4_K_M quantisation, which is the quality-per-gigabyte sweet spot for most people:
| Model class | Fits in 24 GB | Notes |
|---|---|---|
| 7–8B | Comfortably | Very fast; leaves room for a large context window |
| 13–14B | Comfortably | The everyday tier for this machine |
| 32B | Yes | The reason to choose 24 GB; expect a usable conversational speed |
| 70B | No, not usefully | Needs ~40 GB at Q4; a second card or a different machine |
The rough arithmetic, which the VRAM calculator explainer works through properly: a Q4_K_M model occupies roughly 0.6 GB per billion parameters, so 32B lands near 19 GB, leaving a few gigabytes for the KV cache. Long context consumes that remainder quickly, which is why 32B on a 24 GB card is comfortable at ordinary context lengths and tight at very long ones.
Path B’s 16 GB card handles the 13–14B tier the same way and does not reach 32B at Q4.
Before you order: verification checklist
This is the part a specification cannot do for you. Each of these is a number to confirm against the specific parts in your basket, not the category.
- GPU length against case clearance. Find the length of the exact card model — third-party 3090s range from roughly 310 to 340 mm — and check it against the case’s stated maximum with front fans fitted, which is often 20–30 mm less than the headline figure.
- GPU slot width. Many 3090s are 2.7 or 3 slots. Confirm nothing below the card is blocked, and that the case has the slots.
- Cooler height against case width, and cooler width against DIMM height if you are using air.
- PSU connector count and type. A 3090 typically needs three 8-pin PCIe connectors, or a native 12V-2×6 depending on the model. Confirm which your card has and that your supply provides it without adapters where possible.
- Motherboard BIOS version against CPU support. Boards that have sat in a warehouse ship old firmware. Confirm the board supports your CPU out of the box, or that it has USB BIOS flashback so you can update without a working CPU.
- Memory on the board’s QVL at the speed you intend to run, in a two-DIMM configuration.
- Front-panel and case fan headers. Trivial and routinely forgotten.
- If buying a used card: ask for photographs of the PCB and the power connectors, confirm the fans spin freely, and budget for replacing the thermal pads. The 3090’s GDDR6X modules sit on the back of the board and ran hot from new; degraded pads are the normal condition of a five-year-old card, not a defect unique to a bad one.
If any of those checks fails, the failure is in this specification’s applicability to your parts, not in your reading of it. Substitute and continue.
Where to spend more, and in what order
If the budget stretches, the order is not the one most build guides give.
- More VRAM. Always first. A 24 GB card over a 16 GB card changes what the machine can do; nothing else on this list does.
- A second NVMe drive. The cheapest quality-of-life improvement, and the one you will want soonest.
- More system memory, if you work with several large models in a session.
- A better case, with better fans. Sustained loads reward this more than the spec sheet suggests.
- A faster CPU. Genuinely last. It will not make the machine generate tokens faster.
The upgrade that is not on this list is a second graphics card. Splitting a model across two cards works and is how people reach 70B-class models at home, but it brings power delivery, physical spacing and motherboard lane allocation problems that are better solved by planning for them at build time than by retrofitting. That is a different architecture, not an upgrade to this one — and it is specified separately in the dual-GPU 70B machine. If you expect to get there in stages rather than at once, the staged-upgrade PC covers what to over-specify now so the option survives.
Frequently asked questions
Have you actually built this machine?
No, and that is why the page is called a reference architecture rather than a build guide. The components are specified from published manufacturer documentation and are compatible on paper. Physical fitment in your specific case, with your specific card, is what the verification checklist exists to catch — it is the part that cannot be established from a spec sheet.
Is a used RTX 3090 a sensible thing to buy in 2026?
For local AI, yes, with eyes open. Twenty-four gigabytes at 936 GB/s is still a strong specification for language models, and capacity has aged far more slowly than compute. You are accepting no warranty, higher power draw, and a card whose memory modules ran hot from new — budget for replacing the thermal pads and ask the seller for photographs before committing.
Why only 8 CPU cores at this budget?
Because during generation the CPU is idle. Once the model is resident in VRAM the graphics card does the work and the CPU feeds it. Spending the difference on cores rather than VRAM produces a machine that runs smaller models — which is a straightforward downgrade for this workload, however good it looks on a spec sheet.
Do I really need an 850 W power supply for a 350 W card?
For this card, effectively yes. GDDR6X-era NVIDIA cards draw brief current spikes well above rated board power, and those transients are what trip over-current protection — a machine that reboots under load rather than crashing is nearly always an undersized or older supply. The ATX 3.1 standard in the specification matters as much as the wattage, because it defines tolerance for exactly that behaviour.
Can I use an Intel CPU instead?
Yes. Nothing in this architecture depends on AMD — the requirement is a modern platform with PCIe 5.0, DDR5 and enough single-thread speed to keep the card fed. AM5 is specified because it currently offers that at lower power and cost, and because the socket has a longer stated upgrade life. Substituting an equivalent Intel platform changes the board and memory choices, not the reasoning.
Will this run a 70B model?
Not usefully. A 70B model at Q4 needs roughly 40 GB before any context, so it will not fit in either card. It can be made to run by offloading layers to system memory, at speeds slow enough that most people stop doing it within a week. Reaching that tier at home means two cards or a unified-memory machine, both of which are different architectures rather than upgrades to this one — [the dual-GPU 70B machine](/builds/dual-gpu-70b-machine/) covers the first.
Should I buy a prebuilt instead?
If you do not want to do the verification work, genuinely yes — and there is no shame in that answer. Prebuilt machines arrive assembled, tested and under a single warranty, and the guide to [AI workstations](/workstations/best-ai-workstations/) covers them. Building wins on cost per gigabyte of VRAM and on knowing exactly what is inside; it costs you an afternoon and the checklist above.