LLM.co — Private AI & custom AI developmentCall +1 (206) 844-1326All inference local

Securing a Self-Hosted LLM Behind a Firewall

A concrete network and access-control blueprint for locking down a self-hosted LLM behind the enterprise firewall, from egress rules to SIEM and SSO.

Nate Nead8 min read
A reinforced server rack enclosed by concentric defensive walls at dusk

Most private LLM breaches do not start with a model exploit. They start with a port left open, a service account with a static key, or a model file pulled from a public repository without a checksum. The infrastructure beneath the model is boring plumbing, and boring plumbing is where the leaks happen.

Consider the base rate. Censys found more than 10,600 exposed Ollama instances across 1,229 different networks, with no authentication in front of them. That is the easy failure mode: self-hosting does not mean private by default. Private means segmented, authenticated, logged, and monitored, with the model treated as a sensitive asset in its own right.

This piece is the configuration-level reference. It assumes you already decided to run open-weight models inside your own perimeter and now need to hand a network engineer something concrete.

Zone the Inference Stack Like a Payment Network

Flat networks are the first thing to go. Treat the inference stack the way a bank treats cardholder data: separate zones, documented flows between them, and deny-by-default between every pair.

A workable layout has five zones. A user zone where clients and the SSO identity provider live. A DMZ with the reverse proxy and the API gateway. An application zone running the orchestration layer, retrieval service, and agent runtimes. An inference zone with the GPU hosts and the model server (vLLM, TGI, or similar). A data zone holding the vector database, document store, and audit sinks.

Traffic flows one way: user to DMZ, DMZ to app, app to inference, app to data. The inference zone never initiates outbound traffic to the user zone. The GPU hosts should not be able to reach the public internet at all once bootstrapped. Enforce zone boundaries with a stateful firewall that logs denied flows, and repeat the same policy at the host level with nftables or Cilium network policies so a misconfigured switch does not become a bypass.

The zero-trust default is a good anchor here. NIST SP 800-207 is worth citing in the architecture document because auditors recognize it, not because it will tell your network engineer anything new.

Blast Radius by Network Zone
Blast Radius by Network ZoneUser zone (endpoints): 25; DMZ (reverse proxy): 45; App zone (orchestrator): 70; Inference zone (GPU hosts): 90; Data zone (vectors, WORM logs): 10025User zone(endpoints)45DMZ (reverse proxy)70App zone(orchestrator)90Inference zone (GPUhosts)100Data zone (vectors,WORM logs)
Illustrative: relative blast radius if a single workload in each zone is compromised, assuming the segmentation controls described above. Illustrative: a visual comparison, not measured data.

Lock Egress Before You Worry About Ingress

Ingress is the obvious attack surface. Egress is the one that actually leaks data. An inference host with unrestricted outbound DNS and HTTPS can exfiltrate a corpus one chunked request at a time, and nothing in your front-door WAF will see it.

The default on every GPU host and every container in the app zone should be egress deny. From there, build a short allowlist:

  • Your internal package mirror (Artifactory, Nexus, or a pull-through proxy for PyPI and Hugging Face).
  • Your internal container registry.
  • Your model weight store, usually an S3-compatible bucket in the data zone.
  • Your telemetry and SIEM endpoints.
  • Your identity provider's OIDC endpoints, if the service validates tokens directly.

Nothing else. No huggingface.co, no github.com, no pypi.org from production. Mirror everything through a proxy that logs the artifact hash and the requesting workload identity. For agentic workloads that genuinely need to call external APIs, route them through a dedicated egress gateway with a per-tool allowlist and per-request logging. The pattern matters more for agents than for base inference, which is why private AI agents deserve their own egress tier rather than inheriting the model server's.

Terminate TLS Once, Then mTLS Everywhere Inside

The reverse proxy pattern that survives audit is simple. A single ingress tier (Envoy, HAProxy, or NGINX) terminates public TLS, enforces OIDC through your identity provider, strips cookies, and forwards to the API gateway over mutual TLS. Every hop after that is mTLS with short-lived certificates issued by an internal CA, usually SPIFFE/SPIRE or a Vault PKI engine. Rotate every 24 hours. Pin SPIFFE IDs in the service mesh so a stolen cert from one workload cannot impersonate another.

At the API gateway, do three things that are easy to forget. Enforce request size limits (a 50MB prompt is almost always an attack or a bug). Rate-limit per authenticated principal, not per IP, since every call looks like it came from the gateway. Strip and re-emit a correlation ID that follows the request through retrieval, inference, and audit, so a single incident ticket can be reconstructed later.

A wax seal on an envelope resting across a GPU card

Treat Model Weights as Signed Binaries

Model files are executable in every way that matters. JFrog documented roughly 100 malicious model files on Hugging Face in 2024 using Python's pickle format to embed reverse shells that fire when the model is loaded. A separate three-month study surfaced 91 malicious models and 9 malicious dataset loading scripts, with 76 relying on pickle deserialization and 15 on Keras Lambda layer abuse. And the "safe" path turned out not to be safe either: CVE-2025-32434 showed that torch.load(weights_only=True) could still reach arbitrary code execution before the 2.6.0 patch.

The operational response is a short checklist your deployment pipeline enforces, not a policy document:

  • Prefer the safetensors format over pickle-based .bin or .pt files. Reject the latter at ingest.
  • Pin a SHA-256 for every weight file in a signed manifest. The model server refuses to load anything whose hash is not in the manifest.
  • Sign the manifest itself with Sigstore/Cosign and verify the signature on the GPU host at startup.
  • Scan incoming weights with ModelScan or Protect AI's equivalent before promotion to the production bucket.
  • Keep the model registry in your own object store. Treat Hugging Face as an upstream mirror, never as a runtime dependency.

Country of origin is a governance decision, not a technical one, but it is worth noting that Chinese open-weight models now account for 46.4% of routed token usage on OpenRouter, versus 35.7% for US-origin models. If your compliance posture requires provenance review, the review has to happen at ingest, not after the weights are already resident on a GPU. The same reasoning applies when choosing what to download in the first place, which is why we keep a vetted set at open-source LLM downloads.

Malicious Model Files Found on Hugging Face (3-month study)
Malicious Model Files Found on Hugging Face (3-month study)Pickle deserialization exploits: 76; Keras Lambda layer abuse: 15; Malicious dataset loading scripts: 976%15%9%Pickle deserialization exploits76 · 76%Keras Lambda layer abuse15 · 15%Malicious dataset loading scripts9 · 9.0%
Illustrative: a visual comparison, not measured data.

Wire SSO and SCIM Before the First User Logs In

Authentication and provisioning are where self-hosted stacks accumulate shadow access. The pattern that holds up: OIDC at the proxy for interactive users, SCIM 2.0 from the identity provider (Okta, Entra, Ping) for group lifecycle, and short-lived JWTs scoped per request. Workloads use SPIFFE identities, not API keys.

Map identity-provider groups to three roles at minimum. Readers can submit prompts and see their own history. Operators can see system metrics and manage model loadouts. Auditors have read-only access to the full prompt and response log, including retrieved document IDs, but cannot submit prompts. Service accounts that previously used long-lived keys get rotated to workload identities with a 1-hour TTL, and the old keys are revoked at the gateway, not just deprecated.

Agent frameworks deserve extra scrutiny. A recent study on agent-based attack vectors found every tested LLM could be compromised through inter-agent trust exploitation, with context-dependent behaviors creating exploitable blind spots. The practical implication: an agent should authenticate to every tool it calls using its own identity, never a shared service principal, and the audit log should record which agent instance made which tool call.

Pipe Everything to the SIEM, Including the Prompts

The audit requirement most teams underestimate is prompt and response logging. HIPAA, GLBA, and SOC 2 all treat the LLM like any other application that touches regulated data, which means a durable, tamper-evident record of inputs and outputs. Dropping that into Splunk, Elastic, Chronicle, or Sentinel through the usual Fluent Bit or Vector pipeline is the easy part. Deciding what to redact before it leaves the inference zone is the hard part.

A workable sink layout:

  • Structured application logs (JSON) from the gateway, orchestrator, and model server. Include correlation ID, principal, model name, token counts, latency, and tool calls. Not raw prompts.
  • A separate, encrypted stream for prompt and response payloads, keyed to the correlation ID and written to a WORM-configured bucket. Access is auditor-role only, with every read logged.
  • Network flow logs from the firewall and service mesh, tagged by zone.
  • Host telemetry (osquery, auditd, or EDR) from GPU hosts, with alerts on any outbound connection outside the allowlist and any new process under the model server UID.

Build detections for the LLM-specific cases. Repeated prompts followed by chunked low-entropy responses may indicate exfiltration via the model. A sudden spike in tool-call volume from a single agent identity may indicate prompt injection. A model load event not preceded by a signature verification is a direct policy violation. The OWASP Top 10 for LLM Applications is a useful mapping document when writing the correlation rules.

Research backs the compliance pressure. Kong's 2025 Enterprise AI report puts data privacy and security as the top adoption barrier for 44% of organizations. The SIEM wiring is where that concern either gets answered or gets waved away. Teams that want a deeper walkthrough of the audit evidence an assessor will actually ask for can start with our notes on SOC 2, HIPAA, and GDPR, and we cover the same ground operationally in an LLM security audit.

Control Effort vs Audit Value
Control Effort vs Audit ValueEgress allowlist: 30; Weight signature verification: 40; mTLS everywhere: 70; Prompt/response WORM log: 55; SCIM provisioning: 35; Flat network with host firewalls: 80 → →Egress allowlistWeight signature…mTLS everywherePrompt/response W…SCIM provisioningFlat network with…
Illustrative: where to spend engineering hours when the auditor arrives. Higher is better on both axes. Illustrative: a visual comparison, not measured data.

Rehearse the Failure Modes

The controls above do not survive first contact unless someone tries to break them. Schedule a quarterly exercise that includes a model-poisoning drill (introduce a known-bad pickle to the ingest pipeline and verify it is rejected), an egress drill (attempt a DNS-tunneled exfiltration from a GPU host and verify the SIEM alerts), and an identity drill (revoke an operator's IdP account and confirm active sessions terminate within the token TTL). The findings from the first run are usually unflattering and always worth more than the architecture diagram.

Private AI is not a product you install. It is a running system with the same operational discipline you apply to the rest of the regulated estate. The firewall rules, the manifest signatures, and the SIEM pipelines are what let you say "the data stays home" and show the receipts when someone asks.

// written by
Nate Nead
CEO, Nead, LLC

Nate Nead is the CEO of Nead, LLC, a specialized AI development and private AI implementation firm. He and his team help organizations design, build and deploy private AI systems, including on-premise and private-cloud language models, retrieval pipelines and AI agents, that keep sensitive data under the client's control. Over more than 15 years, Nate has founded and grown more than a dozen technology and services brands, and has worked with teams at some of the most well-recognized companies in the Fortune 1000.

Bringing AI in-house, the right way.

Talk through your private or on-prem LLM deployment with an expert who has shipped them in regulated environments.

// the briefing

Private AI, in your inbox.

Occasional, high-signal notes on enterprise LLM deployment, security, and model strategy. No spam.