What a Private GPU Cluster for 200 Users Actually Costs
A line-item breakdown of what an on-prem AI cluster for 200 knowledge workers costs over three years, and which decisions move the number most.

Most CIOs assume a private GPU cluster is a seven-figure experiment that only pays back at the scale of a hyperscaler. In practice, a 200-user deployment for a regulated firm sits in a narrower, more defensible band once you itemize the pieces honestly. The number moves on four decisions: which GPU, how dense the facility can run, how much headroom the model needs, and whether the workload is isolated from the public internet.
The model below is a three-year total cost of ownership for a private LLM infrastructure serving roughly 200 named users at a mid-sized bank, hospital network, or defense contractor. It assumes concurrent utilization of about 15 to 25 users at peak, a mix of chat, retrieval-augmented generation, and document analysis against a 13B to 70B open-weight model, and a Tier 3 facility the client already operates or leases. Numbers are current to late 2026 and reflect US pricing.
So what does the invoice actually look like.
Sizing the Cluster Before You Price It
Two hundred users is not two hundred concurrent sessions. For internal productivity workloads, planners typically assume 10 to 15 percent concurrency at peak, which puts the serving target at roughly 20 to 30 live requests, each wanting a few hundred output tokens at interactive latency. That demand profile fits comfortably inside a single 8-GPU node for most 13B to 70B models at FP8, with a second node for high availability and burst.
GPU choice drives everything downstream. The H100 is the production default for sub-70B serving on a mature stack, while the L40S is the cost-per-token winner for smaller models at moderate concurrency. L40S at roughly $0.72 per hour versus H100 SXM5 at $3.84 per hour on comparable spot markets illustrates the gap, and H100 only overtakes on cost efficiency once you push into high-batch 70B inference or NVLink-dependent training. For 200 users on a 13B model with retrieval, two L40S nodes will serve the load. For 200 users on a 70B model with long-context document analysis, you want H100s.
Headroom matters because a cluster at 90 percent utilization has no capacity for the fine-tune job someone will inevitably request. Build to 60 percent steady-state and you have room to grow without a procurement cycle.
- 11Count named vs concurrent users · 10-15% peak concurrency · 20-30 live requests
- 22Pick model size and precision · 13B-70B at FP8 or FP16 · 70B FP8 or 13B FP16
- 33Estimate tokens per request · Chat + RAG + doc analysis · A few hundred output tokens
- 44Choose GPU class · L40S for small, H100 for 70B · H100 for long-context 70B
- 55Set utilization headroom · Build to 60% steady-state · Leaves room for fine-tunes
- 66Add HA node · Second 8-GPU node for burst · Two nodes total
The Hardware Line Items
A defensible build for this user count is two 8-GPU nodes plus a management and storage tier. For an H100 configuration, a DGX H100 system packs 8 SXM5 GPUs with 640GB of HBM3 and draws about 10.2 kW. Complete 8-GPU DGX H100 systems run $250,000 to $400,000 outright, with OEM HGX H100 equivalents from Supermicro or Dell landing in the same band. Two nodes puts GPU hardware at $500,000 to $800,000.
Networking for a two-node cluster is modest but not optional. Plan on a pair of 400G InfiniBand or RoCE switches for east-west GPU traffic, a 100G Ethernet aggregation pair for storage and management, and transceivers. Budget $80,000 to $150,000 depending on redundancy and whether existing fabric can be reused.
Storage splits into two tiers. A hot tier of NVMe for model weights, KV cache spill, and vector indexes wants roughly 100 to 200 TB usable on an all-flash array; $150,000 to $300,000 is reasonable for a resilient configuration from Pure, VAST, or WEKA. A warm tier for training data, logs, and backups adds another $50,000 to $100,000 on bulk NVMe or hybrid. Call storage $200,000 to $400,000 total.
Management and head nodes, out-of-band management, PDUs, and racks round out capital expenditure at $40,000 to $80,000. That brings three-year capex to a range of roughly $820,000 to $1.43 million for a two-node H100 build. An L40S equivalent that handles smaller models comes in at roughly half.

Power, Cooling, and the PUE Tax
GPU hardware does not run in isolation. Semianalysis estimates the all-in power per H100 SXM chip at about 1.37 kW once host CPU, NICs, and cooling share are included, roughly double the 700W TDP of the chip itself. Two 8-GPU nodes at that load draw about 22 kW continuously, before PUE.
PUE is the multiplier that converts IT load into facility load. The Uptime Institute's 2024 Global Data Center Survey put the industry average PUE at 1.56, essentially flat over several years. Hyperscalers do dramatically better, with Google reporting fleet-wide PUE of 1.09 and Meta 1.08 in 2024, but a mid-sized enterprise facility is realistically between 1.4 and 1.6. At 22 kW IT load and PUE 1.5, total facility draw is 33 kW.
If the cluster lives in colocation, that density sits in high-density AI territory. High-density AI racks drawing 20 kW or more run $3,000 to $6,000 per month before bandwidth, and purpose-built AI colocation for 20 to 40 kW cabinets lands between $3,900 and $8,000 per month. For two cabinets over three years, colocation and power run $280,000 to $580,000. On-premises in an existing enterprise data center, you pay utility rates and amortized facility cost, often lower but with capex for cooling upgrades if the room was not designed for direct-to-chip liquid.
Software, Support, and the People Who Keep It Running
Open-weight stacks are free to download and expensive to operate well. Plan on vLLM or TensorRT-LLM for serving, a vector database (Qdrant, Milvus, or pgvector at this scale), an orchestration layer, observability, and identity integration. If you buy commercial support for the model runtime, Red Hat OpenShift AI, or NVIDIA AI Enterprise, add roughly $4,500 per GPU per year for the NVIDIA stack. That is $70,000 to $100,000 annually across 16 GPUs, or $210,000 to $300,000 over three years.
Staffing is where TCO estimates most often understate reality. A production private LLM infrastructure needs at least one dedicated platform engineer, with partial allocation from a security engineer, an ML engineer for model work, and a data engineer for retrieval pipelines. MLOps engineer base pay runs $90,000 to $257,000 in 2026, with national averages between $130,000 and $165,000. Fully loaded with benefits, a 1.5 to 2.0 FTE allocation lands at $400,000 to $700,000 over three years.
Vendor hardware support (next-business-day or four-hour parts) typically runs 8 to 12 percent of hardware list annually. For an $800,000 build, that is $200,000 to $300,000 over three years.
What Swings the Total
Four decisions move the three-year number by hundreds of thousands of dollars in either direction.
- Quantization. Serving a 70B model in FP8 instead of FP16 roughly halves VRAM and often lets you drop from eight GPUs to four for the same throughput. On a 70B workload this can shift the right answer from H100 to L40S entirely.
- Concurrency assumptions. If the real peak is 50 live requests rather than 20, you need a third node and the networking to connect it, adding $300,000 to $500,000 in capex plus power.
- High availability. Active-active across two sites doubles hardware and roughly doubles facility cost. Active-passive with cold standby is cheaper but accepts a recovery window.
- Air-gap. A fully disconnected network for CMMC Level 2, classified work, or hospital-side PHI isolation adds procurement friction, manual patch pipelines, and typically 15 to 25 percent to staffing cost over the life of the system. The architecture is well understood, but the operational tempo is slower by design.
For regulated firms, the air-gap decision is often not optional. An IBM survey found 57 percent of enterprises cite data privacy as the biggest inhibitor to AI adoption, and for firms under HIPAA, GLBA, or CMMC the inhibitor is a hard legal boundary rather than a preference. LLM.co's work on air-gapped private AI deployment covers the patterns that keep inference inside the perimeter without sacrificing update velocity entirely.
The Three-Year Number
Rolling the line items together, a two-node H100 cluster serving 200 users with high availability lands at roughly $1.9 million to $3.2 million over three years, inclusive of hardware, networking, storage, colocation or on-prem facility share, software support, and staffing. An L40S-based build for the same user count on smaller models runs roughly $1.2 million to $1.9 million. These numbers align with public estimates: three-year TCO for a single DGX H100 is estimated at $480,000 to $750,000 once power, cooling, support, and admin labor are included, and a 200-user production system is two nodes plus the supporting tier.
Whether that number is defensible depends on what the firm would otherwise spend sending the same workload to a hosted API, and on what the compliance team values about keeping the data in the building. For most firms we work with through our custom AI development services, the arithmetic favors on-prem once sustained token volume crosses the point where API bills exceed roughly $40,000 to $60,000 per month. Below that, hybrid is usually the right answer. Above it, the TCO case tightens quickly, and the non-financial benefits of a private, operator-controlled stack start to carry real weight.
The itemization matters more than the headline. A build quoted at $1.5 million that forgot staffing is a $2.3 million build with a surprise. If you want a line-item review of your own workload against this model, our team is happy to walk through it; the engagement process starts with a sizing conversation, not a sales deck.
Throughout his extensive 15+ year journey as a digital marketer, Sam has left an indelible mark on both small businesses and Fortune 500 enterprises alike. His portfolio boasts collaborations with esteemed entities such as NASDAQ OMX, eBay, Duncan Hines, Drew Barrymore, Price Benowitz LLP, a prominent law firm based in Washington, DC, and the esteemed human rights organization Amnesty International. In his role as a technical SEO and digital marketing strategist, Sam takes the helm of all paid and organic operations teams, steering client SEO services, link building initiatives, and white label digital marketing partnerships to unparalleled success. An esteemed thought leader in the industry, Sam is a recurring speaker at the esteemed Search Marketing Expo conference series and has graced the TEDx stage with his insights. Today, he channels his expertise into direct collaboration with high-end clients spanning diverse verticals, where he meticulously crafts strategies to optimize on and off-site SEO ROI through the seamless integration of content marketing and link building.
Bringing AI in-house, the right way.
Talk through your private or on-prem LLM deployment with an expert who has shipped them in regulated environments.
Private AI, in your inbox.
Occasional, high-signal notes on enterprise LLM deployment, security, and model strategy. No spam.














