The Multi-Tenant GPU Machine: A Reference Architecture
The multi-user server assumes its users are colleagues. Remove that assumption and almost everything changes — consumer cards stop qualifying, and prefix caching becomes a leak.
Servers, racks, power and the supporting infrastructure that keeps a self-hosted lab running reliably.
The multi-user server assumes its users are colleagues. Remove that assumption and almost everything changes — consumer cards stop qualifying, and prefix caching becomes a leak.
A machine nobody can physically touch is specified around its recovery ladder. Out-of-band management, network-bound unlock, watchdogs and the runbook a stranger will read under pressure.
Serving many people inverts the arithmetic: batching makes concurrency nearly free, and the KV cache — not the model — is what fills the memory. Specified against latency targets.
A machine that is switched on while you sleep is bought on a different axis entirely. Idle draw, keep-alive, cold-start latency and the API you must not expose — specified.
A NAS for model weights is optimised differently from one for media — throughput, cache and memory matter, and the transcoding engine everyone advertises does not.
Homelab nodes need I/O, cores, low idle power and working IOMMU — not memory bandwidth. Which mini PCs deliver that, what idle power really costs, and where people overspend.
As an Amazon Associate, AI Gear Stack earns from qualifying purchases. Amazon and the Amazon logo are trademarks of Amazon.com, Inc. or its affiliates.