There is no single best AI workstation, and any guide that names one is answering a question nobody asked.
What there is, is a fork in the road. One path buys you speed: a discrete GPU with fast memory, which runs whatever fits in 24 to 32 GB several times quicker than anything else at the price. The other buys you capacity: a unified-memory machine with 128 GB or more, which runs models the first path cannot load at all, considerably more slowly.
Almost every workstation decision reduces to which of those two you need. This guide is organised around that.
Our picks at a glance
-
Best Overall
NVIDIA GeForce RTX 5090
The fastest interactive local AI you can build, for models up to 32B
-
Best for Local LLMs
Framework Desktop (Ryzen AI Max+ 395)
128 GB unified memory reaches 70B on x86, at around 120 W
-
Best for Developers
NVIDIA DGX Spark
Develop against 70B-class models on the same CUDA stack you deploy to
-
Best Premium
Apple Mac Studio (M4 Max)
546 GB/s against 128 GB — the best large-model speed at this power level
-
Best High-End Option
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
96 GB and full discrete bandwidth in one card, with no compromise
-
Most VRAM per Dollar
Apple Mac Studio (M3 Ultra)
Up to 512 GB — the only desktop that runs frontier-scale open models
Which side of the fork are you on?
Answer one question: what is the largest model you need to run?
If the answer is 32B parameters or smaller, you want a discrete GPU. A 32B model at 4-bit needs about 23 GB including context, which fits comfortably in a 32 GB card, and it will generate at 60 to 75 tokens per second — faster than you can read.
If the answer is 70B or larger, no consumer graphics card will hold it. A 70B model at 4-bit needs roughly 48 GB. Your options are two GPUs, a workstation card, or unified memory. The reasoning behind these figures is in How Much VRAM Do You Need for Local AI?
The trap is buying for the model you aspire to run rather than the one you will actually use. A 70B model at six tokens per second sounds impressive and is genuinely tedious to work with. A 32B model at seventy tokens per second is a tool you will reach for daily.
The specifications
| Product | VRAM | Bandwidth | Runs up to | Power | Local LLMs | Where to buy |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 NVIDIA | 32 GB | 1792 GB/s | 32B at Q4 entirely in VRAM, with room for long context | 575 W total board power; 1,000 W system PSU recommended | Excellent | Check Price on Amazon NVIDIA GeForce RTX 5090 at Amazon — opens in a new tab |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition NVIDIA | 96 GB | 1792 GB/s | 70B at Q8, or 120B-class at Q4, entirely in VRAM | 600 W (Max-Q variant is configurable to 300 W) | Excellent | Check Price on Amazon NVIDIA RTX PRO 6000 Blackwell Workstation Edition at Amazon — opens in a new tab |
| Framework Desktop (Ryzen AI Max+ 395) Framework | 96 GB | 256 GB/s | 70B at Q4, using the large GPU memory allocation | Approximately 120 W typical under sustained inference load | Excellent | Check Price on Amazon Framework Desktop (Ryzen AI Max+ 395) at Amazon — opens in a new tab |
| NVIDIA DGX Spark NVIDIA | 128 GB | 273 GB/s | 70B at Q4 on one unit; roughly 200B-class across two linked units | Approximately 240 W from a single USB-C power supply | Excellent | Check Price on Amazon NVIDIA DGX Spark at Amazon — opens in a new tab |
| Apple Mac Studio (M4 Max) Apple | 128 GB | 546 GB/s | 70B at Q4 with substantial context headroom | Well under 200 W under sustained load | Excellent | Check Price on Amazon Apple Mac Studio (M4 Max) at Amazon — opens in a new tab |
| Apple Mac Studio (M3 Ultra) Apple | 512 GB | 819 GB/s | 400B-class at Q4, or 70B at full FP16 precision | — | Excellent | Check Price on Amazon Apple Mac Studio (M3 Ultra) at Amazon — opens in a new tab |
Read the bandwidth column against the capacity column and the whole trade-off is visible in one view. The DGX Spark and the Framework Desktop hold four times the memory of an RTX 5090 at roughly a seventh of the bandwidth.
The recommendations
Best overall for most people
Best Overall
NVIDIA GeForce RTX 5090
4.1/5AI Gear Stack review score: 4.1 out of 5
Best for The fastest single-card local inference you can buy without moving to workstation pricing
The 32 GB frame buffer and 1.8 TB/s of bandwidth make this the only consumer card that runs 32B-class models comfortably in VRAM. The cost is real: 575 W, a 1,000 W PSU and a case that can move the heat.
- VRAM
- 32 GB
- Bandwidth
- 1792 GB/s
- GPU
- GB202 Blackwell, 21,760 CUDA cores
- Runs up to
- 32B at Q4 entirely in VRAM, with room for long context
Strengths
- 32 GB VRAM clears the 30B-class barrier that 24 GB cards sit just under
- 1,792 GB/s is roughly 78% more bandwidth than an RTX 4090, and bandwidth is what sets token speed
- Native FP4 support meaningfully increases effective capacity on supported runtimes
- CUDA remains the path of least resistance for every local AI toolchain
Trade-offs
- 575 W board power demands a serious PSU and case airflow
- Still short of the 48 GB needed to hold a 70B model at Q4
- Physically large; verify clearance before buying
- Availability and street pricing have been erratic since launch
For the majority of people building a workstation for local AI, a machine built around an RTX 5090 is the right answer.
32 GB holds a 32B-class model at 4-bit with real context headroom, and 1,792 GB/s runs it fast enough that generation stays ahead of reading. That combination — a genuinely capable model at genuinely interactive speed — is what makes local AI feel like a tool rather than a demonstration.
Build around it accordingly: a 1,000 W power supply, a case that can actually exhaust 575 W, and 32 GB of system RAM. The CPU matters far less than people assume; once weights are resident on the GPU, a mid-range processor is not the constraint.
Best for large models
Best for Local LLMs
Framework Desktop (Ryzen AI Max+ 395)
Best for Running 70B-class models locally on x86 without a 600 W power budget
The most practical x86 route to 70B-class local inference. Memory capacity is the thing that decides what you can run at all, and 128 GB of it at 256 GB/s beats any consumer discrete card on capacity by a wide margin.
- VRAM
- 96 GB
- Memory
- 128 GB
- Bandwidth
- 256 GB/s
- GPU
- Radeon 8060S, 40 RDNA 3.5 compute units
Strengths
- Up to 96 GB addressable by the GPU — far beyond any consumer discrete card
- Standard x86, so every tool works without architecture caveats
- Mini-ITX and roughly 120 W under load
- Framework's repairability and parts availability
Trade-offs
- 256 GB/s is a seventh of an RTX 5090's bandwidth
- Memory is soldered — the configuration you buy is the one you keep
- ROCm rather than CUDA, with the ecosystem gaps that implies
If your requirement genuinely is 70B, this is the most practical x86 route to it.
Up to 96 GB addressable by the GPU means a 70B model at 4-bit fits with room for context, in a 4.5-litre chassis drawing around 120 W. Reaching the same capacity with graphics cards means two of them, roughly 900 W, and a build designed around thermals.
The honest limitation is speed. At 256 GB/s, a 70B model generates at roughly five to six tokens per second. That is fine for batch work, for overnight processing, and for tasks where you send a prompt and come back. It is frustrating for anything conversational.
Memory is soldered, so the configuration you buy is the one you keep. Buy 128 GB.
Best for developing against large models
Best for Developers
NVIDIA DGX Spark
Best for Developing against 70B-class models on the same CUDA stack you will deploy to
128 GB of CUDA-addressable memory in a box that draws about 240 W. It loads models a 5090 cannot touch, then runs them at roughly a seventh of the bandwidth. That trade is excellent for development and poor for serving.
- VRAM
- 128 GB
- Memory
- 128 GB
- Bandwidth
- 273 GB/s
- GPU
- Blackwell GPU with 5th-generation Tensor Cores
Strengths
- 128 GB unified memory holds a 70B model at Q4 with context to spare
- Full CUDA stack — the same code runs on datacentre Blackwell
- 240 W and near-silent for the capability on offer
- Two units link over 200 GbE to pool memory for larger models
Trade-offs
- 273 GB/s means single-digit to low-double-digit tokens per second on 70B
- Arm64 Linux; a few tools still assume x86
- Priced well above a comparable x86 machine with a discrete GPU
The DGX Spark is a different proposition from everything else here: it is a development machine for people who will deploy to NVIDIA hardware.
128 GB of CUDA-addressable memory means you can develop and test against 70B-class models using exactly the same stack — CUDA, the same libraries, the same kernels — that runs on datacentre Blackwell. That parity is the entire point, and it is worth real money to teams whose production target is NVIDIA.
At 273 GB/s it is not a serving machine. Two units can be linked over 200 GbE to pool memory for larger models, which is a genuinely useful capability at this size and power draw.
Best if you are not tied to CUDA
Best Premium
Apple Mac Studio (M4 Max)
Best for Large models at usable speed, silently, on a desk
546 GB/s against 128 GB of unified memory is the best capacity-and-bandwidth combination available at this power level. MLX and llama.cpp are both well optimised for it; CUDA-only tooling is not an option.
- VRAM
- 128 GB
- Memory
- 128 GB
- Bandwidth
- 546 GB/s
- GPU
- 40-core Apple GPU with hardware ray tracing
Strengths
- 546 GB/s is roughly double any Strix Halo machine
- 128 GB unified memory runs 70B at Q4 comfortably
- Near-silent and remarkably power-efficient
- MLX and llama.cpp are both mature on Apple silicon
Trade-offs
- No CUDA; a real constraint for image generation and fine-tuning
- Memory cannot be upgraded after purchase
- Apple's memory pricing is steep
546 GB/s against 128 GB of unified memory is the best capacity-and-bandwidth combination available at this power level, by a wide margin.
A 70B model at 4-bit runs at roughly 13 tokens per second — more than double what any LPDDR5X-based x86 machine manages, and enough to be usable interactively. It does so near-silently and under 200 W.
The constraint is CUDA, or rather its absence. llama.cpp and Apple’s MLX are both mature and fast on Apple silicon, so language models are well served. Image generation and fine-tuning workflows frequently assume CUDA, and support ranges from good to nonexistent depending on the tool.
The largest models on a desk
Most VRAM per Dollar
Apple Mac Studio (M3 Ultra)
Best for The largest models anyone can run on a single desktop machine
Up to 512 GB of unified memory at 819 GB/s. Nothing else on a desk runs models this large at this speed, and the price reflects exactly that.
- VRAM
- 512 GB
- Memory
- 512 GB
- Bandwidth
- 819 GB/s
- GPU
- 80-core Apple GPU
Strengths
- 512 GB of addressable memory is unmatched in a desktop
- 819 GB/s bandwidth, roughly triple a Strix Halo machine
- Runs frontier-scale open models no consumer GPU can load
Trade-offs
- Extremely expensive in the large-memory configurations
- No CUDA
- M3-generation cores in an M4-generation lineup
Up to 512 GB of unified memory at 819 GB/s. Nothing else on a desk comes close, and the configurations that matter are extremely expensive.
It exists on this list because it is the only machine that runs frontier-scale open models — 400B-class at 4-bit — without a rack. If that is the requirement, this is the answer, and the price is what it is.
Best single-card workstation
Best High-End Option
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
4.2/5AI Gear Stack review score: 4.2 out of 5
Best for Running or fine-tuning 70B-class models on a single card
96 GB of ECC GDDR7 at RTX 5090 bandwidth. It is the single-card answer to 70B-class work, priced as a professional tool rather than a consumer component.
- VRAM
- 96 GB
- Bandwidth
- 1792 GB/s
- GPU
- GB202 Blackwell, 24,064 CUDA cores
- Runs up to
- 70B at Q8, or 120B-class at Q4, entirely in VRAM
Strengths
- 96 GB holds a 70B model at Q8 with context to spare
- Same 1,792 GB/s bandwidth as the RTX 5090
- ECC memory and the professional driver branch for long unattended runs
- Max-Q variant fits 300 W power envelopes
Trade-offs
- Professional pricing, several times an RTX 5090
- 600 W in the full-power configuration
- Overkill for anything that fits in 32 GB
96 GB of ECC GDDR7 at 1,792 GB/s: 70B-class capacity and discrete-GPU bandwidth, in one card.
This is the configuration that has no compromise in it, and it is priced accordingly. For fine-tuning in particular — where memory pressure is far higher than inference and long unattended runs benefit from ECC — it is the option that actually works rather than the one that nearly works.
Building around the choice
Memory
For a discrete-GPU machine, 32 GB of system RAM is the sensible default. Once weights are on the GPU, system memory handles loading and the operating system. 64 GB is worth it if you expect to offload large models or run several services alongside. We work through where the cutover falls in 64GB vs 128GB RAM for Local AI.
For a unified-memory machine, system RAM is model memory. Buy the maximum you can afford, because it cannot be changed later.
Storage
Capacity over speed. Models are large and accumulate quickly — 1 TB is a sensible floor if you intend to experiment. Sequential throughput affects load time only; it does nothing for generation. See Best NVMe SSDs for AI Workloads.
Power and cooling
This is where AI workstation builds most often go wrong. A 575 W graphics card under a sustained inference load is a fundamentally different thermal problem from the same card gaming: the load is constant rather than bursty, and it can run for hours.
Budget a 1,000 W supply for an RTX 5090, verify physical clearance before ordering the case, and plan the airflow properly. A machine that thermally throttles after twenty minutes is slower than a lesser one that does not.
CPU
Less important than almost any other component for inference. A modern mid-range processor with enough PCIe lanes for your GPU is sufficient. Spend the difference on VRAM — it is the specification that determines what you can run at all.
Common questions
Is a Mac good for AI work?
For language models, yes — unified memory means a 128 GB Mac Studio runs models that would need a workstation GPU on a PC, and at higher bandwidth than any x86 unified-memory machine. For image generation and fine-tuning, the absence of CUDA is a real constraint that no amount of memory compensates for.
Should I buy a prebuilt or build my own?
For a discrete-GPU machine, building gives you control over the power supply, cooling and case — all of which matter more under sustained inference load than under gaming load. Unified-memory machines are not buildable; they are bought as configured.
How much system RAM do I need alongside a GPU?
32 GB is the sensible default. Model weights live on the GPU, so system memory handles loading and the operating system. Go to 64 GB if you expect to offload large models or run several services at once.
Can I add a second GPU later?
Often, but plan for it up front. You need the PCIe slots, the physical clearance, and a power supply sized for both cards. Retrofitting a second 450 W card into a build that was not designed for it usually means a new supply and frequently a new case.
Is a workstation GPU worth it over a consumer card?
Only if you need the capacity. The RTX PRO 6000 costs several times an RTX 5090 and delivers the same bandwidth — what you are buying is 96 GB instead of 32 GB, plus ECC. If your models fit in 32 GB, it buys you nothing.
Continue your research
- Best GPUs for Local AI — if you have settled on the discrete-GPU path
- Best Mini PCs for Local LLMs — the unified-memory path in detail
- How Much VRAM Do You Need for Local AI? — size your requirement first
- 64GB vs 128GB RAM for Local AI — how much system memory is worth buying