A reference architecture is a component specification with the reasoning attached: every part chosen for what it contributes to running models locally, every trade-off stated, and every compatibility requirement written down as something you can check.
What these are, and what they are not
These are specifications, not tested builds. Each one is derived from published manufacturer documentation — socket, TDP, connector type, memory support, physical dimensions — and the parts are specified to work together on paper.
AI Gear Stack has not assembled these machines. Nobody here has put the cooler on the board, closed the case, or measured whether the graphics card clears the drive cage in the specific chassis you buy. That matters, because the things that break a build are rarely the things on a spec sheet: a heatsink that fouls the first memory slot, a card that is four millimetres too long for a case the manufacturer lists as supporting it, a revision of a board that ships with older firmware than the CPU requires.
So each architecture specifies components by requirement first and by example second — “an 850 W ATX 3.1 unit with a native 12V-2×6 connector” is the specification; a named model is an illustration of one. That is deliberate. It is the form of advice that stays true when a specific SKU is discontinued, and it is the honest limit of what can be established without a bench.
Every architecture ends with a verification checklist: the specific measurements and part numbers to confirm against your own basket before you place the order.
What each architecture includes
- A complete component specification: CPU, GPU, motherboard, memory, storage, power supply, cooling and case
- Why each part is there, and what we considered instead
- What the finished machine can actually run, with quantisation stated
- Where to spend more, and where not to — on an AI build the answer is almost always “more VRAM, cheaper everything else”
- Power delivery and physical clearance figures, which is where most builds go wrong
- A pre-order verification checklist, because a specification you have not checked against your own parts is a hypothesis
Why we publish these rather than tested builds
Assembling and testing every configuration would produce stronger claims, and one day it should. Until then the choice is between publishing nothing and publishing the component reasoning with its basis stated plainly.
The reasoning is the part that transfers. Which tier of card to buy, why memory capacity beats memory speed for this workload, why the CPU matters far less than it does for gaming, how much power headroom a transient-spiking GPU actually needs — none of that changes when a model number does, and all of it is establishable from documentation. What is not establishable that way is fitment in your particular case, and we say so at the point where it matters rather than in a footnote.
If you want recommendations for complete machines that ship assembled, the workstation guides cover those instead.
Published architectures
- Build The Machine You Already Own: A Reference Architecture Every other build here starts from a blank sheet. This one starts from an inventory — find the one component actually stopping…
- Build The Multi-Tenant GPU Machine: A Reference Architecture The multi-user server assumes its users are colleagues. Remove that assumption and almost everything changes — consumer cards stop qualifying, and prefix…
- Build The Reproducible AI Machine: A Reference Architecture Same model, same prompt, same seed, same GPU — and still a different answer. Why that happens, when it matters, and the…
- Build The Dual-Purpose AI and Gaming PC: A Reference Architecture This site tells you to ignore gaming rankings. If you also game, you cannot — the two workloads want almost disjoint parts…
- Build The Staged-Upgrade AI PC: A Reference Architecture You cannot future-proof compute, only the chassis around it. A machine bought over a year or two — what to over-specify at…
- Build The Colocated AI Machine: A Reference Architecture A machine nobody can physically touch is specified around its recovery ladder. Out-of-band management, network-bound unlock, watchdogs and the runbook a stranger…
- Build The Air-Gapped AI Machine: A Reference Architecture Most organisations that say air-gapped need no data egress instead. For the ones that genuinely need isolation: provisioning, verification, patch cadence and…
- Build The Multi-User Inference Server: A Reference Architecture Serving many people inverts the arithmetic: batching makes concurrency nearly free, and the KV cache — not the model — is what…
- Build The Fine-Tuning Workstation: A Reference Architecture Training inverts almost everything this site says about buying hardware — and most people should rent instead. The memory arithmetic, the inversions,…
- Build The Always-On AI Server: A Reference Architecture A machine that is switched on while you sleep is bought on a different axis entirely. Idle draw, keep-alive, cold-start latency and…
- Build The Dual-GPU 70B Machine: A Reference Architecture Two used 24 GB cards is the cheapest route to 48 GB of CUDA memory — and the point where a home…
- Build The $2,000 Local AI PC: A Reference Architecture A component specification for a machine that runs 32B-class models well — with the reasoning behind every part, and an explicit checklist…
Component guides to start from
- Review AMD Radeon RX 9070 XT Review: A Better Card That Runs Models Worse Newer, more efficient and architecturally improved — with 50% less memory and 49% less bandwidth than the card it follows. For local…
- Review AMD Radeon RX 7900 XTX Review: The Hardware Is Not the Problem The only new card near this price with 24 GB, and it beats both 16 GB NVIDIA options on the specifications that…
- Review NVIDIA RTX PRO 6000 Blackwell Review: Not Faster, Just Able The same memory bandwidth as an RTX 5090 with three times the capacity — so you are not buying speed. You are…
- Review NVIDIA GeForce RTX 4060 Ti 16GB Review: Superseded, and Worth It Only on Price The slowest card reviewed here, replaced by one with 56% more bandwidth for 15 W more — and the single card on…
- Review NVIDIA GeForce RTX 3090 Review: Still the Answer, and Not Forever Eleven articles on this site lean on this card. Twenty-four gigabytes for the least money is still unbeaten — and the risks…
- Review NVIDIA GeForce RTX 5060 Ti 16GB Review: The Cheapest Way Into 16 GB, and the Slowest Exactly half an RTX 5070 Ti's bandwidth at exactly the same capacity — and the only way to get 16 GB of…