Where the Money Actually Goes in a Private AI Deployment

A line-item financial breakdown of private, on-prem LLM deployments, from GPUs and power to RAG pipelines, security audits, and staffing.

Nate Nead8 min read
High-density GPU server racks with liquid cooling in a modern enterprise data center

Private AI pitches tend to collapse into one number: the GPU server. That number is real, and it is large, but it accounts for less than half of a three-year private LLM budget. The rest hides in power contracts, networking refreshes, vector database licensing, red-team retainers, and the salary line for the people who keep the thing running. CIOs who approve a project on hardware economics alone end up explaining variance to the board in year two.

This piece walks the full line-item breakdown of a reference on-prem deployment sized for a mid-market regulated enterprise: a single 8-GPU inference node, a smaller training/fine-tuning cluster, a RAG stack, and the governance scaffolding that keeps auditors satisfied. Every figure below is drawn from vendor pricing, salary benchmarks, or published engineering benchmarks. The ranges are wide because real deployments are wide. The point is to give a finance team a defensible shape for the budget, not to pretend there is one true number.

Compute Hardware Is the Anchor, Not the Total

The compute line dominates year-one capex. A single 8-GPU HGX H100 server runs $250,000 to $320,000 fully built, with per-GPU cost landing near $33,000 once NVSwitch fabric, NICs, and integration are folded in. PCIe variants shave 10 to 15 percent off that figure but give up NVLink, which matters for larger model inference and almost all fine-tuning.

Most private LLM programs we see at LLM.co's custom deployment practice start with one inference node and a smaller training rig, not a 32-GPU cluster. A reasonable reference stack is one 8xH100 inference server, one 4xH100 or 4xH200 fine-tuning server, head nodes, and spares. That lands total compute capex between $450,000 and $750,000 before anything is racked. Teams weighing a self-built cluster against a turnkey system should read the appliance versus DIY comparison before committing, because the integration labor is not trivial.

Refresh is the quieter half of this line. GPU generations are landing roughly every 18 to 24 months, and resale values collapse the moment the next tier ships. Budget a 36-month amortization and plan for partial refresh in year three, not a full forklift.

Year-One Hardware Capex vs. Three-Year Run Cost
Year-One Hardware Capex vs. Three-Year Run Cost8xH100 inference node: 285,000; 4xH100 fine-tune node: 180,000; Networking fabric: 140,000; Storage tiers: 180,000; Power + cooling: 0Year 1 capex3-year run cost8xH100 inference node285,000180,0004xH100 fine-tune node180,000120,000Networking fabric140,00060,000Storage tiers180,00090,000Power + cooling0150,000
Illustrative: a visual comparison, not measured data.

Power and Cooling Are a Standing Cost Center

Power is where on-prem economics meet physics. A single H100 server draws around 10 kW and costs roughly $10,500 to $10,700 per year in power at $0.12/kWh average US commercial rates, with cooling adding another 25 to 40 percent on top. Two nodes plus head gear plus storage push a typical pilot into the $35,000 to $50,000 annual utility line.

The deeper problem is rack density. A conventional data center rack draws 7 to 10 kW, but an AI-capable rack can demand 30 kW to over 100 kW, averaging more than 60 kW in dedicated AI facilities. Legacy colos were not built for this. Expect quotes for high-density cage upgrades, liquid cooling retrofits, or a premium per-kW tariff from a facility that already handles AI workloads. In corporate data centers, the server load itself is only part of the total: servers account for roughly 60 percent of data center electricity, rising toward 75 percent in hyperscaler AI facilities, with the balance going to cooling, power distribution, and lighting.

Close-up of an 8-GPU AI server chassis showing heatsinks and cooling plates

Networking and Storage Rarely Get Costed Correctly

Multi-GPU training requires low-latency fabric. InfiniBand NDR switches, cables, and NICs for an 8-to-16 GPU footprint typically add $75,000 to $200,000 depending on whether the deployment runs NVLink-only within a node or scales across nodes. Ethernet-based alternatives using 400GbE RoCE are closing the gap, but the switch count and cabling cost remain substantive.

Storage splits into three tiers. Hot NVMe for model weights and KV caches, warm all-flash for RAG corpora and embeddings, and cold object storage for training data archives and audit logs. For a reference deployment: 50 to 100 TB NVMe at roughly $400 to $600 per TB, 200 to 500 TB flash in the mid-single-digit dollars per TB-month if leased, and object storage priced per GB. All-in storage capex plus three-year operating cost typically runs $150,000 to $400,000. The vector database is a separate line, and the case for owning that layer rather than renting it has gotten stronger as pricing on managed services has drifted up.

Models, Fine-Tuning, and the Inference Runtime

Open-weights models from Llama, Mistral, Qwen, and DeepSeek carry no per-token licensing fee, which is the main reason finance teams justify private deployment in the first place. The cost sits in fine-tuning compute and the engineering around it. A single supervised fine-tuning run on a 70B-class model consumes between $15,000 and $80,000 in GPU-hours when done on owned hardware, substantially more when rented. Teams running quarterly refreshes should budget $100,000 to $300,000 annually in training compute alone.

Commercial model licensing exists as a hybrid option. Some enterprises run an open-weights model for high-volume routine traffic and route sensitive or complex requests to a licensed closed model deployed in VPC. Pricing for enterprise licenses varies by vendor and is rarely published, but assume six-figure annual minimums for any serious volume. For the pure open-source path, the economics compared to API consumption are covered in detail in how private LLMs replace API subscriptions.

The inference runtime itself (vLLM, TensorRT-LLM, SGLang, or a managed equivalent) is mostly open source, but security-hardened variants and vendor support contracts for production serving typically run $50,000 to $200,000 per year.

Where a Three-Year Private LLM Budget Goes
Where a Three-Year Private LLM Budget GoesStaff and engineering: 48%; Hardware and refresh: 22%; Power, cooling, networking: 10%; RAG, agents, runtime: 9%; Security, audits, compliance: 7%; Fine-tuning compute: 4%Staff and engineeri…48%Hardware and refresh22%Power, cooling, net…10%RAG, agents, runtime9%Security, audits, c…7%Fine-tuning compute4%
Illustrative: a visual comparison, not measured data.

RAG, Agents, and the Application Layer

The retrieval and orchestration layer is where most pilot budgets quietly double. A production RAG stack includes a vector database (Qdrant, Weaviate, Milvus, or pgvector at scale), an embedding model with its own GPU footprint, a reranker, document ingestion pipelines, and the glue code that turns all of it into something the business can call. Expect $80,000 to $250,000 in year-one build cost, with ongoing engineering to maintain connectors as source systems change.

Agentic workloads add another layer: tool execution environments, sandboxes, policy engines, and traffic monitoring. The reasons this needs to run inside your perimeter rather than on a public endpoint are laid out in the piece on on-prem isolation for autonomous agents, and the operational tooling for it is the subject of the private AI agents practice. Confidential computing on H100 and H200 adds measurable but tolerable overhead: less than 7 percent on H100 inference in one benchmark, and under 9 percent average across H100 and H200 TEE tests, with a 20 to 25 percent increase in time-to-first-token due to attestation.

Security, Audits, and Compliance Overhead

Compliance is a recurring line, not a project cost. A SOC 2 Type 2 audit for a mid-market enterprise with AI in scope runs $30,000 to $100,000+ for larger organizations, repeating annually. Penetration testing adds $15,000 to $50,000 per year. HIPAA or GDPR-aligned program work, if not already in place, adds consulting fees that typically land between $50,000 and $200,000 in year one.

LLM-specific audit work is a separate category. Red teaming against prompt injection, data exfiltration via tool use, jailbreak chains, and training-data poisoning is now standard for regulated deployments. A full LLM security audit engagement typically runs $40,000 to $150,000, and most buyers in finance and healthcare now repeat it annually. Model governance tooling and documentation workflows add roughly $25,000 to $75,000 per year in SaaS and labor combined.

Deployment Path by Cost Center
Cost centerSelf-built on-premManaged appliancePrivate cloud VPC
Hardware capex$450K-$750KBundled lease$0 capex, higher opex
Power and cooling$35K-$50K/yrIncluded or passthroughIncluded in instance price
MLOps staffing3-5 FTE1-2 FTE2-3 FTE
Compliance scopeFull ownershipShared with vendorShared with cloud
Refresh cadenceCustomer ownsVendor-managedInstance-level
Typical 3-yr TCO$7.8M-$12.1M$5M-$9M$6M-$14M
Illustrative: a visual comparison, not measured data.

Staff Is the Largest Line Nobody Puts on the Slide

A functioning private LLM program needs a platform lead, two or three ML or LLM engineers, an MLOps or SRE for the GPU fleet, a data engineer for the RAG layer, and fractional security and compliance support. US machine learning engineers earn a median total compensation of $278,000 per Levels.fyi, with specialists in generative AI and fine-tuning commanding 40 to 60 percent premiums above baseline ML salaries. Loaded cost for a five-person team lands between $1.8M and $2.8M annually, before any security or compliance headcount.

This is why staffing almost always exceeds hardware over a three-year horizon, and why so many programs fail organizationally before they fail technically. The structural problem is covered in the AI center of excellence guide, and the operational lesson is simple: hardware depreciates on a schedule, but salaries compound.

What a Three-Year Reference Budget Actually Looks Like

Rolling it up for a mid-market regulated deployment with one 8xH100 inference node, one 4xH100 fine-tuning node, a production RAG stack, SOC 2 and either HIPAA or GDPR alignment, and a five-person team:

  • Year 1: $2.8M to $4.5M, weighted toward hardware capex, build labor, and initial compliance.
  • Year 2: $2.4M to $3.6M, mostly staff, power, training compute, and recurring audits.
  • Year 3: $2.6M to $4.0M, with partial hardware refresh and expanded fine-tuning cadence.

Three-year TCO lands between $7.8M and $12.1M. For comparison, continuous operation of a Llama-3 70B on an AWS p4d.24xlarge runs roughly $287,000 per year per instance before any data egress, support, or compliance costs, and most enterprises need multiple instances. The break-even against a serious public API spend typically arrives in year two for high-volume workloads, later for low-volume ones. The decision calculus is laid out further in the RFP requirements checklist.

How to Defend the Budget Line by Line

A CFO challenge almost never kills the hardware line. It kills the undefended ones: the vector database bill that showed up in month four, the SOC 2 scope expansion in month ten, the second MLOps hire in month fourteen. Treat every line in this breakdown as a standing cost center with an owner, a renewal date, and a documented alternative. Price the refresh before year one closes. Price the audit cadence before the first auditor arrives. Price the staff ramp against a hiring market where generative AI specialists are the scarcest line item in the plan.

Private AI is not expensive because the GPUs are expensive. It is expensive because it is a platform, and platforms carry recurring cost. The organizations that make it work treat it as such from the first budget cycle, and defend each line with the same rigor they apply to the rest of the infrastructure estate.

// written by
Nate Nead
CEO, Nead, LLC

Nate Nead is the CEO of Nead, LLC, a specialized AI development and private AI implementation firm. He and his team help organizations design, build and deploy private AI systems, including on-premise and private-cloud language models, retrieval pipelines and AI agents, that keep sensitive data under the client's control. Over more than 15 years, Nate has founded and grown more than a dozen technology and services brands, and has worked with teams at some of the most well-recognized companies in the Fortune 1000.

Bringing AI in-house, the right way.

Talk through your private or on-prem LLM deployment with an expert who has shipped them in regulated environments.

// the briefing

Private AI, in your inbox.

Occasional, high-signal notes on enterprise LLM deployment, security, and model strategy. No spam.