The verdict
4.2 /5
The same bandwidth as an RTX 5090 with three times the memory, so you are not buying speed — you are buying 70B at Q8, ECC and MIG, none of which exist below it. Two 5090s do not reach it: 64 GB does not hold that model, at nearly twice the power. Overkill for anything that fits in 32 GB, and the only answer for several things that do not.
Scored against our methodology
Buy it if
Fine-tuning 70B-class models, serving mutually untrusted tenants, work that must be reproducible, or running 70B at Q8 in a single machine — cases where nothing below it qualifies at all rather than merely more slowly.
Skip it if
Anything that fits in 32 GB, where an RTX 5090 is exactly as fast for dramatically less; homelab use, where 600 W and professional pricing make no sense; or reaching 70B at Q4 cheaply, which two used 24 GB cards do.
Strengths
- 96 GB runs 70B at Q8 entirely in VRAM — no consumer card does this at any price
- 1,792 GB/s, matching the fastest consumer card
- ECC memory, which makes long runs trustworthy rather than merely possible
- MIG partitioning, which GeForce cards cannot do at all
- No datacenter licence ambiguity for commercial deployment
- Max-Q variant delivers the same capacity at 300 W
- 96 GB is also what makes partitioning useful rather than theoretical
Trade-offs
- Identical bandwidth to an RTX 5090 — no faster at all on anything that fits in 32 GB
- The premium over a consumer flagship is very large
- 600 W in standard form
- Professional driver branch is not optimised for games
- Not a homelab part: power, price and feature set are aimed at organisations
- Two used 24 GB cards reach 70B at Q4 for a fraction of the price
Nine articles on this site reach for this card, six of them build architectures. It is the answer in the fine-tuning workstation, the multi-tenant machine, the air-gapped machine and three more — and until now there was no review behind any of that.
The reason it keeps appearing is not that it is fast. It has the same memory bandwidth as an RTX 5090. What it has instead is three capabilities that do not exist anywhere below it, and each one is why a different architecture names it.
| Identity | |
|---|---|
| Manufacturer | NVIDIA |
| Product family | RTX PRO Blackwell |
| Model | RTX PRO 6000 Blackwell |
| Form factor | Graphics card |
| Release year | 2025 |
| Price class | Flagship ($3,500+) |
| Compute | |
| GPU | GB202 Blackwell, 24,064 CUDA cores |
| GPU architecture | Blackwell (GB202) |
| VRAM (GB) | 96 |
| VRAM type | GDDR7 ECC, 512-bit |
| Memory bandwidth (GB/s) | 1792 |
| Compute capability note | ECC memory, MIG partitioning, professional driver branch |
| Connectivity & physical | |
| Power draw | 600 W (Max-Q variant is configurable to 300 W) |
| Workload suitability | |
| Local LLMs | Excellent |
| Ollama | Excellent |
| Stable Diffusion / ComfyUI | Excellent |
| Fine-tuning | Excellent |
| Homelab | Limited |
| Developer workstation | Excellent |
| Largest comfortable model | 70B at Q8, or 120B-class at Q4, entirely in VRAM |
You are not buying speed
Set the two cards side by side and the shape of the product is immediately clear.
| RTX 5090 | RTX PRO 6000 Blackwell | |
|---|---|---|
| Memory | 32 GB GDDR7 | 96 GB GDDR7 ECC |
| Memory bandwidth | 1,792 GB/s | 1,792 GB/s |
| CUDA cores | 21,760 | 24,064 |
| Board power | 575 W | 600 W |
| Partitioning | None | MIG |
Identical bandwidth. Three times the memory.
Since capacity decides what runs and bandwidth decides how fast, that means something specific: for any model that fits in 32 GB, these two cards generate at the same speed and one costs a great deal more. The 5090 review is the better purchase for that work and this review will not pretend otherwise.
What the extra 64 GB buys is not a faster answer. It is an answer at all.
| Model | Roughly | RTX 5090 (32 GB) | PRO 6000 (96 GB) |
|---|---|---|---|
| 32B at Q4 | ~19 GB | Yes | Yes |
| 70B at Q4 | ~42 GB | No | Yes, comfortably |
| 70B at Q8 | ~75 GB | No | Yes |
| 120B-class at Q4 | ~72 GB | No | Yes |
That 70B-at-Q8 row is the one that has no substitute. Running a 70B model at eight-bit rather than four-bit quantisation, entirely in VRAM, on one card, is not something any consumer part does at any price.
The comparison people expect to win, and it does not
The obvious objection is that two consumer flagships should be cheaper than one professional card. It is worth working through, because the arithmetic does not go where people assume.
| 2 × RTX 5090 | 1 × PRO 6000 | |
|---|---|---|
| Total memory | 64 GB | 96 GB |
| Board power | 1,150 W | 600 W |
| Slots occupied | Two, spaced | One |
| Interconnect | PCIe only — no NVLink on Blackwell GeForce | Not needed |
| Software | Tensor parallelism required for speed | Nothing to split |
| 70B at Q8 | No | Yes |
Two 5090s do not reach where one of these goes. 64 GB does not hold a 70B model at Q8, so the workaround fails at the specific job the card exists for — and it fails while drawing nearly twice the power, occupying two slots, and requiring the tensor-parallel setup the dual-GPU architecture spends most of its length on.
The multi-card route is genuinely the right answer one tier down, with a pair of 24 GB cards reaching 70B at Q4 for far less money. It stops being the answer here.
The three things that do not exist below it
Each of these is why a different build on this site names this card specifically, and none of them is available on any GeForce part at any price.
ECC memory
A single flipped bit during inference produces one odd token and nobody notices. During hour nine of a twelve-hour fine-tuning run it corrupts a gradient, which corrupts the weights, which propagates through every step after it — and you find out at evaluation, if at all. The fine-tuning workstation argues this at length.
The reproducible machine makes the sharper version: a flipped bit produces an answer you can neither reproduce nor explain, in a context where explanation is the entire point.
MIG partitioning
Hardware partitioning, with separate memory and fault domains per instance. GeForce cards cannot do this at all — not slowly, not with software, not at all — which is the sharpest hardware constraint on this site and the reason the multi-tenant machine rules consumer cards out entirely rather than merely discouraging them.
96 GB is also what makes partitioning useful rather than theoretical. Divide 32 GB four ways and each tenant gets 8 GB; divide 96 GB four ways and each gets something that runs a 32B model.
No datacenter licence question
NVIDIA’s GeForce driver licence has historically restricted datacenter deployment, and whether an office machine serving colleagues falls under that is genuinely unclear — the multi-user server states the ambiguity rather than resolving it, because we are not lawyers.
The professional line is the product sold for that use. Part of what the premium buys is not having the conversation, and for a commercial deployment that has real value even though it appears on no specification sheet.
Max-Q, and the 300 W version of this card
The Max-Q variant is configurable to 300 W against the standard card’s 600, and it deserves more attention than it usually gets.
Halving the power budget costs real throughput, but far less than half — and it changes what the card can go in. A 300 W professional card with 96 GB fits thermal and power envelopes that a 600 W one does not: a rack unit with a fixed circuit allocation, a workstation without a 1,000 W supply, a machine that runs continuously where the always-on server arithmetic applies.
If the capacity is what you need and the last few percent of throughput is not, this is the more sensible card of the two.
What it is not
Not a gaming card, and not sold as one. The professional driver branch prioritises stability and application certification over per-title optimisation. It will run games perfectly well and you are paying a great deal for silicon whose value is elsewhere.
Not for anything that fits in 32 GB. The product record for this card says so plainly and it is worth repeating: identical bandwidth to a 5090 means identical speed on any model both can hold. If your work is 32B-class models, buy the consumer card.
Not a homelab part. 600 W, professional pricing and a feature set built around ECC and partitioning are aimed at organisations, not enthusiasts. The always-on server reaches for a 180 W card for a reason.
Not the cheapest route to 70B. Two used 24 GB cards reach 70B at Q4 for a fraction of the price. What they do not reach is Q8, or a single-slot deployment, or ECC.
Who should buy one
Buy it if you are fine-tuning 70B-class models, serving mutually untrusted tenants, doing work that must be reproducible, or running 70B at Q8 in one machine. In each case nothing below it qualifies — not more slowly, but at all.
Buy the Max-Q variant if the capacity is what you need and 300 W fits your enclosure or circuit better than 600 W does. For most of the roles above, it is the more sensible card.
Buy an RTX 5090 instead if your models fit in 32 GB. Same bandwidth, same speed on that work, dramatically less money.
Buy two used RTX 3090s instead if you want 70B at Q4 for the least money and can accept the tensor-parallel setup, 700 W of graphics cards and no warranty. That is what the dual-GPU architecture is for.
Frequently asked questions
Is the RTX PRO 6000 Blackwell worth it over an RTX 5090?
Only if you need what the extra memory unlocks. Both cards have 1,792 GB/s of bandwidth, so on any model that fits in 32 GB they generate at the same speed and the 5090 costs dramatically less. The professional card earns its premium at 70B — comfortably at Q4, and at Q8 entirely in VRAM, which no consumer part manages at any price — plus ECC, MIG partitioning and the absence of a datacenter licence question.
Would two RTX 5090s be cheaper and better?
Cheaper, and not better for the job this card exists for. Two of them give 64 GB, which does not hold a 70B model at Q8 — so the workaround fails at the specific target, while drawing about 1,150 W across two slots and requiring tensor parallelism to use the bandwidth. Blackwell GeForce cards have no NVLink either. Multi-card is genuinely the right answer one tier down, with two 24 GB cards reaching 70B at Q4 cheaply.
What can it run that a consumer card cannot?
70B at Q8 entirely in VRAM, at roughly 75 GB, and 120B-class models at Q4 around 72 GB. A 32 GB card reaches 32B at Q4 and stops. The gap is not speed — the two cards read memory at the same rate — it is that one of them can hold the model and the other cannot, and capacity is a threshold rather than a gradient.
Why does ECC matter here when consumer cards manage without it?
Because the consequences scale with how long the work runs and how much it is trusted. A flipped bit during a chat produces one odd token nobody notices. During hour nine of a twelve-hour fine-tuning run it corrupts a gradient, then the weights, then every step after — discovered at evaluation if at all. On a machine whose output must be reproducible it is worse still: an answer you can neither repeat nor explain, in a context where explanation is the entire point.
What is MIG and do I need it?
Multi-Instance GPU partitions the hardware itself, giving each instance its own memory and fault domain. You need it when several people who should not read each other’s data share one card — GeForce parts cannot do it at all, so for mutually untrusted tenants a consumer card is not a cheaper option but not an option. If your users are colleagues, you do not need it and the multi-user server architecture covers that case instead.
Should I get the Max-Q variant?
For most of the roles this card suits, yes. It runs the same 96 GB at 300 W rather than 600, which costs real throughput but considerably less than half — and it changes what enclosures and circuits the card fits. A rack unit with a fixed power allocation, a workstation without a 1,000 W supply, or a machine that runs continuously all favour it. Buy the full-power card when the last few percent of throughput genuinely matters.
Is it a good gaming card?
It will run games well and that is not what it is for. The professional driver branch prioritises stability and application certification over per-title optimisation, and you would be paying a large premium for memory and features that games do not use. For a machine that games and runs models, the dual-purpose build covers the trade properly and reaches a very different conclusion.
Did AI Gear Stack test this card?
No. This assessment is drawn from manufacturer specifications and published architectural detail, as stated at the top of the page. Throughput figures are calculated from the memory bandwidth specification rather than measured, and the behaviour of MIG partitioning and ECC under real workloads is exactly the kind of claim that wants adversarial testing rather than a specification sheet. We say so rather than guessing.