What to Put in a Private LLM RFP Before You Sign

The clauses, SLAs, and technical requirements to demand in a private LLM RFP, and which vendor answers should disqualify a bid before procurement.

Nate Nead7 min read
A steel filing cabinet drawer open to reveal a glowing server module resting on a folder, symbolizing weights held inside the buyer's perimeter.

Most enterprise RFPs for large language models are IT procurement templates with an "AI" section stapled to the front. They score demos, count certifications, and rank vendors on features that a probabilistic system will never deliver deterministically. One recent analysis found that a standard IT RFP misses up to 60% of the risk-relevant questions for AI vendors, because it was built for on-premise software and predictable SaaS, not models that route data through third-party APIs and drift week over week.

The gap matters most at signature. Once the master agreement is executed, the buyer inherits whatever the vendor's default topology, log schema, and termination clause happen to be. What follows is a line-item checklist, written for a CIO, CISO, or head of AI who has already approved the initiative and now needs to disqualify the wrong vendor before the ink dries.

Weights, Data, and IP Ownership Are Non-Negotiable

The first section of any private LLM RFP should force the vendor to answer five ownership questions in writing. Any hedge here is disqualifying.

  • Who owns the buyer's training and fine-tuning data? The buyer, in perpetuity, with no license-back to the vendor for other customers.
  • Who owns the resulting fine-tuned weights? The buyer, with perpetual, irrevocable, sublicensable rights.
  • Are weights exportable in a standard format on termination? Yes, in GGUF or an equivalent portable format, within a defined number of business days.
  • Can the vendor use buyer data or gradients to improve models for anyone else? No, including no aggregated telemetry, no "anonymized" derivatives.
  • What happens to weights and data after termination? Certified destruction within 30 days, with a signed attestation.

Vendors who claim ownership of derived weights will leverage that claim at renewal. The RFP should also require a bill of materials for the base model, including the license (Apache 2.0, Llama Community License, MIT, proprietary), because the base license flows through to the fine-tuned artifact. The cost-benefit case for keeping this stack in-house is covered in our companion analysis of open-weights training economics.

Air-Gap and Update Paths Need to Be Diagrammed, Not Described

Any vendor claiming air-gapped or sovereign deployment should submit a network diagram showing every ingress and egress the platform requires during install, inference, and update. Marketing language is not a topology. Require the following as line items:

  • Offline install media, signed with a published key, verified by SHA-256 checksum on arrival.
  • A documented model refresh path that does not require outbound calls to a vendor CDN, license server, or telemetry endpoint.
  • Signed model updates delivered as portable artifacts, with a rollback procedure to the previous version tested during acceptance.
  • A "no phone home" attestation, with any exceptions (license validation, crash telemetry) enumerated and made optional.

The reason to be specific is that many private LLM stacks look air-gapped until an operator tries to apply a security patch. Our field notes on on-prem deployment friction catalog the recurring failure modes. For customers who want the update path handled as a managed artifact rather than as an internal runbook, an LLM appliance collapses this into a signed image cycle with a documented rollback window.

Where On-Prem LLM Spend Actually Lands
Where On-Prem LLM Spend Actually LandsChips and staff (combined): 75; Electricity (3-yr TCO): 15; Cooling overhead: 25; MLOps engineer (annual, USD k): 13575Chips and staff…Electricity (3-…15Cooling overhead25135MLOps engineer…
Chips and staff dominate the cost stack by an order of magnitude; electricity is real but never decisive. Illustrative: a visual comparison, not measured data.

Audit Log Schemas Belong in the Contract, Not the Roadmap

Logging is where most AI RFPs go quiet. A vendor that answers "we log all activity" without specifying the schema is deferring the compliance problem to the buyer. The RFP should require a field-level specification for every log stream, delivered before contract signature. At minimum:

  • Request-level fields: timestamp, principal identity, tenant, model version, prompt hash, retrieval sources cited, tool calls invoked, tokens in and out, latency, safety filter verdicts.
  • Agent-level fields: parent trace ID, step index, tool arguments, tool response hash, cost attribution, isolation boundary crossed.
  • Administrative fields: model deployment events, weight replacements, fine-tuning jobs, RAG index rebuilds, policy changes, and who approved each.
  • Delivery format: structured JSON, streamed to the customer's SIEM over syslog or an equivalent standard, with retention set by the customer.

The last point matters because logs held in the vendor's tenancy are the vendor's evidence, not yours. Field parity between LLM calls and downstream tool invocations is the specific gap most agent platforms carry, and it is what separates a debuggable incident from a fog of disconnected traces. Buyers deploying private AI agents should treat trace correlation as an acceptance criterion, not a nice-to-have.

Red-Team Acceptance Tests Belong in the Statement of Work

A private LLM should not go into production on the strength of a demo. The RFP should specify red-team acceptance tests that the vendor must pass, on the buyer's infrastructure, against the buyer's data, before final payment. Anchor the test plan to a public taxonomy so scope is not negotiable.

The OWASP Top 10 for LLM Applications 2025 puts prompt injection (LLM01) at the top for the second consecutive edition, because models process instructions and data in the same channel and neither RAG nor fine-tuning fully mitigates the risk. The NIST Generative AI Profile (AI 600-1), released July 26, 2024 as a companion to AI RMF 1.0, enumerates 12 risk categories unique to or exacerbated by generative AI, from confabulation and data privacy to CBRN information capabilities.

A padlock and fountain pen resting on a stack of contract pages, representing signature-time protections in a private LLM procurement.

Translate both frameworks into acceptance criteria: a fixed test set of adversarial prompts, a threshold pass rate, a documented triage process for regressions, and a right to re-test after every model or RAG index update. Guidance on running these engagements is collected in our note on red-teaming a private LLM. Coverage weights should be published in the RFP scoring rubric so vendors know where hedged answers cost them points.

Twelve NIST GenAI Risk Categories, Weighted for RFP Scoring
Twelve NIST GenAI Risk Categories, Weighted for RFP ScoringConfabulation: 14; Prompt injection and misuse: 13; Data privacy: 12; Information security: 12; Information integrity: 10; Harmful bias: 9; Intellectual property: 7; Human-AI configuration: 6; Value chain and component integration: 6; Obscene / illicit content: 5; Environmental impact: 3; CBRN information: 3Confabulation14 · 15%Prompt injection and…13 · 14%Data privacy12 · 13%Information security12 · 13%Information integr…10 · 11%Harmful bias9 · 10%Intellectual prope…7 · 7%Human-AI configura…6 · 6%Value chain and…6 · 6%Obscene / illic…5 · 5%
Illustrative weighting. NIST AI 600-1 enumerates 12 categories; the RFP scoring rubric should assign each an explicit share so hedged answers cost measurable points. Illustrative: a visual comparison, not measured data.

Incident Notification Windows Have to Do the Math

Regulated buyers cannot accept the default incident timelines that appear in vendor commercial contracts. HIPAA requires covered entities to notify affected individuals within 60 days of discovering a breach. Many AI vendor templates default to 72-hour or 30-day incident notification to the customer, meaning a covered entity that signs the default terms may have as little as 30 days left on its own clock by the time it hears about the incident.

Coverage is also uneven across the frontier vendors. As of May 2026, Anthropic, OpenAI, Microsoft, Google, and AWS all sign BAAs but only for specific enterprise product tiers, with consumer and self-serve tiers excluded across all five. The RFP should require notification within 24 hours of the vendor's own discovery, a named breach coordinator, and the right for the customer to lead public disclosure. Broader compliance overlap is mapped in our SOC 2, HIPAA, and GDPR reference.

How a 30-Day Vendor Notification Window Consumes the HIPAA Clock
How a 30-Day Vendor Notification Window Consumes the HIPAA ClockDay 0 (vendor discovers): 0; Day 10: 17; Day 20: 33; Day 30 (vendor notifies buyer): 50; Day 40: 67; Day 50: 83; Day 60 (HIPAA deadline): 10000Day 0 (vend…17Day 1033Day 2050Day 30 (ven…67Day 4083Day 50100Day 60 (HIP…
Illustrative. HIPAA gives covered entities 60 days from discovery; a default 30-day vendor notification consumes half of that budget before the buyer knows an incident occurred. Illustrative: a visual comparison, not measured data.

Exit, Escrow, and the Kill Clause

The final section of the RFP should assume the vendor relationship ends. Three provisions carry the weight.

  • Source and weight escrow with a neutral third party, released on defined triggers: acquisition, insolvency, material breach, or failure to remediate a critical CVE within a stated window.
  • Data portability: fine-tuned weights, RAG indexes, evaluation sets, and log archives exported in documented formats at the customer's request, with a fixed per-export SLA.
  • A performance kill clause: quantitative regression thresholds on the buyer's acceptance test suite trigger a right to terminate for cause without penalty, converting the relationship from term-locked to performance-locked.

Escrow answers a question most private LLM contracts never ask: what happens when the vendor is acquired by a hyperscaler or shuts down a product line. Portability answers what happens when the model drifts and a replacement is required. The kill clause answers what happens when neither party wants to litigate but the platform no longer meets spec. For deal-shape context around these clauses, our contract structures analysis covers how fixed-scope, managed appliance, and co-build agreements distribute the same obligations differently.

What Separates a Signable Vendor from a Shortlisted One

A private LLM RFP is not a features exercise. It is a set of contractual defaults the buyer will live with for three to five years, applied to a system that will change under them. Weights ownership, air-gap update mechanics, log schemas, red-team thresholds, notification windows, and exit terms are the six line items where vendor answers separate the serious from the shortlisted. The rest of the document scores who wrote the better proposal. These sections score who can be signed.

// written by
Nate Nead

Nate Nead is the founder and CEO of Marketer, a distinguished digital marketing agency with a focus on enterprise digital consulting and strategy. For over 15 years, Nate and his team have helped service the digital marketing teams of some of the web's most well-recognized brands. As an industry veteran in all things digital, Nate has founded and grown more than a dozen local and national brands through his expertise in digital marketing. Nate and his team have worked with some of the most well-recognized brands on the Fortune 1000, scaling digital initiatives.

Bringing AI in-house, the right way.

Talk through your private or on-prem LLM deployment with an expert who has shipped them in regulated environments.

// the briefing

Private AI, in your inbox.

Occasional, high-signal notes on enterprise LLM deployment, security, and model strategy. No spam.