How to Rip Out a Public LLM API Contract Without Breaking Production
A step by step plan for replacing an OpenAI or Anthropic contract with a private LLM, including cutover sequencing, prompt parity testing, and rollback.

Most teams that leave a public LLM API do it twice. The first attempt is a weekend swap: point the base URL at a self-hosted endpoint, watch quality collapse on the prompts that mattered most, roll back by Monday standup. The second attempt is slower, duller, and the one that holds. It treats the cutover as a software migration with a shadow period, parity gates, and a written exit from the vendor contract, rather than a config change.
This piece is for the team that has already made the decision. The business case is settled. What remains is sequencing: which workloads move first, what evidence you need before you cut traffic, what stays on the public API indefinitely, and how to end the contract without surprise invoices or orphaned data. The goal is a 30/60/90 plan that nobody has to defend in an incident review.
Decide What You Are Actually Replacing
"Migrate off OpenAI" is not a scope. A single enterprise tenant typically runs chat completions, embeddings, a transcription endpoint, two or three fine-tuned models, a vector store the vendor hosts, and a handful of Assistants or Responses objects that encode tool definitions. Each of those is a separate migration with its own parity criteria.
Inventory first. Pull 60 days of API logs and group calls by endpoint, model, token volume, latency SLO, and the business process that triggered them. Flag anything touching regulated data (PHI, PCI, CJIS, ITAR, attorney work product). Flag anything bound to a vendor-specific object such as a stored Assistant, a hosted file, or a fine-tuned checkpoint. Those are the migration's hard parts, because they do not have a drop-in equivalent on an open-weight stack.
The inventory also tells you what to leave alone. A low-volume, non-sensitive summarization job that costs forty dollars a month is not worth the engineering time to move in the first wave. Workloads that touch board materials, patient records, loan files, or source code with export restrictions are. Rank by sensitivity and volume, not by whichever team shouts loudest.
Stand Up a Gateway Before You Stand Up a Model
The single most useful piece of infrastructure in this migration is a routing layer that speaks the OpenAI wire format on the client side and can dispatch to any backend on the server side. LiteLLM Proxy is the common choice; Hugging Face TGI's Messages API covers the same surface for open models. Either lets applications keep calling /v1/chat/completions while you change what sits behind it.
With the gateway in place, three capabilities become cheap: shadow traffic (send the same request to a candidate model and discard the response), header-based routing (send 5% of a cohort to the private endpoint), and per-tenant policy (block regulated workloads from leaving the VPC). This is the strangler fig pattern applied to inference: the old system keeps serving production while you quietly build its replacement around it, route by route, until the old facade can be removed.
Treat the gateway as the audit boundary. Every request and response, with model id, token counts, latency, and routing decision, lands in your SIEM. That log is what satisfies SOC 2 CC7, HIPAA audit controls, and the EU AI Act's logging obligations once you turn on the private backend. If you need a deeper walkthrough of what the stack looks like in production, our custom LLM stack options post covers the pieces.
Run Shadow Mode Until Parity Is Boring
Shadow mode is where migrations are won or lost. For every production call to the public API, the gateway duplicates the request to the candidate open-weight model (Llama 3.3 70B, Qwen 2.5 72B, Mistral Large, whatever your audit cleared) and logs the pair. Users still see the public API's answer. Nothing changes for them.
What you collect in shadow is the evidence base for the cutover decision. For each workload, define three parity thresholds up front:
- Task accuracy on a held-out eval set — exact match for extraction, ROUGE or BERTScore for summarization, pass@1 for code, LLM-as-judge with a second model for open-ended generation. Pick the metric before you look at results.
- Latency at p95, measured end to end from the gateway. A private model that is 200ms slower is usually fine; one that is 2 seconds slower will change user behavior.
- Refusal and format drift — the rate at which the candidate declines a request the incumbent answered, or returns malformed JSON when a tool call was expected.
Expect the first shadow run to be bad. Prompts written against GPT-4 or Claude rarely transfer intact. A PromptBridge benchmark found that a prompt optimized for GPT-5 that scored 99.39% on HumanEval dropped to 68.70% when ported verbatim to Llama 3.1 70B. The fix is not a better model; it is a second pass on the prompt, the system message, and the few-shot examples, measured against the same eval. Budget two engineering weeks per material workload for this tuning. Our debugging hallucinations in open-source models post goes into the specific failure modes to watch for.

Sequence the Cutover in Three Waves
A workable 30/60/90 looks like this. Days 0 to 30: gateway in production, shadow mode on for the top five workloads by volume, private inference cluster accepting internal traffic only. Days 31 to 60: canary 10% of two workloads where shadow parity held for two weeks, then ramp to 50% if error budgets stay green. Days 61 to 90: full cutover on migrated workloads, decommission the public API calls for those routes, keep the vendor contract alive for the long tail.
Canary before blue-green. A canary release exposes a small slice of real users to the new backend and lets you measure drift under production load, which shadow mode cannot fully simulate because the user never sees the candidate's answer. Blue-green is appropriate once the canary has held at 50% for a week, because by then the "green" environment is a known quantity rather than a hope.
Pick first-wave workloads deliberately. Good candidates: internal search over documentation, meeting summarization, code review assistance, classification tasks with clear ground truth. Bad first candidates: anything facing customers, anything with regulatory exposure measured in millions, anything where the current prompt has been iterated on for a year and nobody remembers why it works. Move the easy wins first so the organization gets a visible delivery before the hard ones start. The failure patterns we see most often in the first 90 days of a private rollout are almost all sequencing errors.
Decide What Stays on the Public API
Not everything should move. A mature private deployment usually keeps a thin, metered connection to one public provider for three categories of work: frontier reasoning tasks where the open-weight gap is still real, bursty workloads that would otherwise require idle GPU capacity, and experimental features where the team is still deciding whether the capability belongs in production at all. Route those through the same gateway, with explicit allowlists on which data classes can leave the network.
Reliability is part of the argument for keeping the connection, in both directions. OpenAI's longest documented outage was a roughly four-and-a-half-hour global failure on 11 December 2024 that took down ChatGPT, the API, and Sora together after a telemetry service overwhelmed the Kubernetes control plane. An empirical outage study put ChatGPT at only 88.85% of days fully accessible. A private cluster you operate is not immune to failures either, but the failures are yours to schedule and your gateway can fail over between the two.
Write the Contract Exit, Not Just the Technical One
Technical cutover without contractual cutover produces a 70% invoice that recurs forever. Before any traffic moves, read the current MSA for notice period (usually 30 or 60 days), minimum commits, and data deletion obligations. Enterprise agreements often include a prepaid credit pool or a monthly minimum that does not shrink when your usage does.
Three deliverables earn the exit:
- A written deprecation notice to internal consumers, mirroring the vendor's own practice. OpenAI gave 12 months of notice before the Assistants API sunset on 26 August 2026. Your internal teams deserve the same courtesy, delivered via RFC 9745 Deprecation and Sunset headers on the gateway so client code logs the warning.
- A data-deletion attestation from the vendor covering fine-tuned model weights, uploaded files, and stored Assistants or vector stores. Keep the attestation with your SOC 2 evidence.
- A reduced floor on the renewal, negotiated against the specific workloads you are keeping. Vendors will discount to retain the account rather than lose it entirely; the leverage disappears once the renewal auto-executes.
Vendor lock-in is now the dominant procurement concern in enterprise AI. In one 2026 survey, 94% of IT leaders reported fearing lock-in as AI gets more deeply embedded, and 80% saw costs exceed forecast. The migration you are running is the proof that lock-in was escapable, which is useful ammunition for the next renewal with any vendor. If the organization is building this muscle for the first time, our private LLM RFP checklist is where the next procurement cycle should start.
What a Clean Cutover Looks Like at Day 90
At the end of a disciplined migration, the picture is unglamorous. Three or four workloads are fully on the private cluster, logged to your SIEM, meeting the parity thresholds you set in week two. One or two workloads are still on the public API because the economics or the capability gap justify it, routed through the same gateway and governed by the same policy. The vendor contract is renewed at a lower floor with a documented exit clause. Nobody wrote a press release.
That is the result worth planning for. The teams that treat the switch as a product launch tend to roll back; the ones that treat it as a strangler-fig migration, with shadow traffic, parity gates, canaries, and a written contract exit, tend to be running on their own infrastructure a year later and quietly adding workloads. If you want help scoping the first wave or sizing the cluster behind it, that is the conversation our custom LLM deployment team has every week.
Nate Nead is the CEO of Nead, LLC, a specialized AI development and private AI implementation firm. He and his team help organizations design, build and deploy private AI systems, including on-premise and private-cloud language models, retrieval pipelines and AI agents, that keep sensitive data under the client's control. Over more than 15 years, Nate has founded and grown more than a dozen technology and services brands, and has worked with teams at some of the most well-recognized companies in the Fortune 1000.
Bringing AI in-house, the right way.
Talk through your private or on-prem LLM deployment with an expert who has shipped them in regulated environments.
Private AI, in your inbox.
Occasional, high-signal notes on enterprise LLM deployment, security, and model strategy. No spam.














