Fine Tuning Your Open Weights Model for Private LLMs
A practical walkthrough of fine tuning Llama, Qwen, and Mistral inside your own VPC or air gap, covering LoRA, QLoRA, data curation, eval, and rollout.

Most teams default to fine-tuning when they should start with retrieval, and reach for retrieval when the right answer is actually a small adapter trained on their own data. The three decisions that matter are which technique, which base model, and how to serve the result inside your network. Get them wrong and you spend a quarter on a model that regresses on tasks the base already handled.
This is the method, not the business case. If you need the cost and ROI argument first, start with the companion piece on training an open weights LLM, then come back here for the recipe.
Pick the Technique Before the Model
Four options cover almost every private workload: continued pretraining, full fine-tuning, LoRA, and QLoRA. The choice is driven by what you are trying to teach the model, not by how much compute you have.
Continued pretraining is for domain language the base model never saw: internal product codes, a drug nomenclature, decades of claim narratives. You run next-token prediction over unlabeled corpora. It is the only technique that reliably shifts the model's underlying distribution, and it is also the easiest way to destroy capability you wanted to keep. In one empirical study, continual instruction training dropped BLOOMZ-7.1b's MMLU-SocialScience score from 36.18% to 26.06%, with the forgetting intensifying as model scale grew from 1B to 7B. Mix in general instruction data at roughly 5–15% of tokens, or do not do it at all.
Full fine-tuning updates every parameter against labeled examples. It is the right call when behavior must change in subtle, pervasive ways: a reasoning style, a refusal policy, a structured output format that LoRA cannot stabilize. The cost is linear in parameters and quadratic in headaches.
LoRA freezes the base and trains small low-rank matrices injected into attention and MLP projections. The original paper reported that against full fine-tuning of GPT-3 175B with Adam, LoRA can reduce trainable parameters by roughly 10,000x, meaning about 0.01% of the model is touched. Quality on narrow tasks is usually within a point or two of full fine-tuning.
QLoRA is LoRA on top of a 4-bit quantized base. It introduced 4-bit NormalFloat, Double Quantization that saves ~0.37 bits per parameter, and Paged Optimizers for handling gradient-checkpointing spikes. The practical consequence: a 70B model that would need around 672 GB in full precision drops to about 46 GB in 4-bit. That is the difference between a cluster and a single card.
If you are new to parameter-efficient methods, DoRA is worth watching. It decomposes weights into magnitude and direction and consistently outperforms LoRA on commonsense reasoning, by 2.1% on LLaMA2-7B and 4.4% on LLaMA3-8B at similar parameter budgets. It drops into most PEFT pipelines with one config change.
| Technique | Best for | Trainable params | Quality vs full FT | Forgetting risk |
|---|---|---|---|---|
| Continued pretraining | New domain vocabulary | 100% | N/A (shifts base) | High |
| Full fine-tuning | Deep behavior change | 100% | 100% | Medium |
| LoRA | Narrow task adaptation | ~0.1–1% | ~95–98% | Low |
| QLoRA | Large models, small budget | ~0.1–1% | ~90–95% | Low |
Match the Base Model to the Workload
Open-weight model choice has stopped being a leaderboard question and started being a licensing and footprint question. The median downloaded model on Hugging Face sits at 406M parameters as of 2025, up only slightly from 326M in 2023, even as the mean rose to 20.8B. Small models remain the workhorses, which should inform what you ship, not just what you prototype.
A rough mapping that holds up in practice:
- Llama 3.1 70B for general reasoning, long-context summarization, and anything a committee will compare against GPT-4-class output. Permissive license with named-use restrictions.
- Llama 3.1 8B or Mistral Small for high-volume classification, extraction, routing, and tool-calling agents where latency and cost dominate.
- Qwen2.5 (7B, 14B, 32B, 72B) for multilingual workloads, strong code performance, and when you want a 32B that punches above its weight.
- Gemma 2 (9B, 27B) for workloads with Google-aligned safety tuning and conservative output.
- DeepSeek V2/V3 for math and code; vet the provenance and licensing posture before adopting, since the weights are separable from the hosted service.
The license is a compliance artifact. Have counsel read it before an ML engineer downloads the checkpoint, not after.

Size the GPUs for the Path You Picked
Memory is the gate; FLOPs are the clock. Both matter, but you cannot start a job that will not fit.
For Llama 3.1 70B, Meta's own numbers show training memory of roughly 500 GB for full fine-tuning, 160 GB for LoRA, and 48 GB for QLoRA. That maps cleanly to on-prem hardware: QLoRA on a single H100 80GB or RTX 6000 Ada; LoRA on 2x H100 with NVLink; full fine-tune on 8x H100 with FSDP or DeepSpeed ZeRO-3. The 8B tier is a different world: QLoRA fits in 14 GB on a single RTX 4090.
Three practical notes before anyone requisitions hardware:
- QLoRA's dequantization overhead adds roughly 20–30% to step time versus BF16, which is usually a fine trade for the memory savings.
- Context length is a memory multiplier. Doubling sequence length can double activation memory even with gradient checkpointing.
- Average GPU utilization in enterprise fine-tuning sits well below 50%. Buy for the job you actually run, not the headline benchmark. See the cluster cost breakdown for how this plays out at 200 users.
Shape the Dataset Before the First Epoch
Data decides outcome more than any hyperparameter you will touch. The failure mode is almost always the same: too little data, too much noise, no held-out eval, and a train/eval split that leaks.
Rough starting volumes, derived from vendor guidance that generalizes well. AWS Bedrock recommends a minimum of 200 samples for Nova fine-tuning with 2,000–10,000 recommended, capped at 10,000 combined records for most models. That is a reasonable instruction-tuning floor for a 7B–13B model. For continued pretraining of a 70B on domain text, think in hundreds of millions to low billions of tokens.
Five checks that prevent most regressions:
- Deduplicate aggressively. MinHash or exact-match dedup at document and chunk level. Duplicates inflate loss curves and memorize PII.
- Hold out a true eval set. Split by time, matter, or customer, not by random row. Random splits leak.
- Label the hard cases. A thousand adversarial examples beats fifty thousand easy ones.
- Scrub for PII and privilege. Especially for legal and healthcare corpora; this is a HIPAA and SOC 2 concern long before it is an ML concern.
- Keep a replay buffer. Mix a slice of general instruction data to blunt catastrophic forgetting.
For structured guidance on how source documents become training examples, the companion piece on structuring data for LLM performance covers the extraction side.
Build an Eval Harness That Catches Regressions
Fine-tuning without evaluation is wishful thinking. The harness has three layers.
First, capability floors: run MMLU, GSM8K, HumanEval, and TruthfulQA against the base and the fine-tune. You are not trying to beat the base; you are trying to not break it. A regression of more than two or three points on a capability you need is a stop-ship.
Second, task evals: a frozen set of your own prompts with graded answers. For generative tasks, pair an LLM-as-judge with a small human-reviewed sample to calibrate drift.
Third, safety and policy evals: refusal behavior, jailbreak resistance, prompt injection, and whatever your regulator cares about. Teams running under the NIST AI RMF should map these to the generative AI profile's recommended controls.
Merge, Serve, and Keep the Adapters Separable
For single-tenant deployments, merge the LoRA weights into the base and serve the result on vLLM or TGI. For multi-tenant or multi-task deployments, do not merge. vLLM supports batching requests with different adapters in the same forward pass, which turns five underutilized GPUs into one shared GPU in the right conditions.
A workable vLLM launch for multi-adapter serving uses --enable-lora, --max-loras, and --max-lora-rank with adapters mounted from a shared volume. Hot-loading adapters at runtime is now production-grade in current releases. Hold adapters in object storage inside your VPC, version them like code, and sign them. Keep an audit log of which adapter answered which request; your compliance team will ask.
Network placement matters as much as the serving flags. Put the inference endpoint behind an authenticating gateway with request logging, rate limits, and content filters. The patterns in the piece on firewalling a self-hosted LLM apply unchanged to a fine-tuned deployment, as do the operational traps covered in what breaks in the first 90 days.
For teams building this inside a regulated environment, the custom deployment practice at LLM.co handles the whole chain, from base selection through adapter versioning and air-gapped serving.
A Monday Morning Decision
If one person on the ML team has to start Monday, the default path is QLoRA on a Llama 3.1 8B or Qwen2.5 14B base, 2,000 to 10,000 cleaned instruction examples, a held-out eval of a few hundred domain prompts, and vLLM behind an internal gateway. That gets you a measurable improvement on a specific workflow in two to four weeks, on hardware that already exists in most data centers. Scale up to 70B QLoRA only when the smaller model's ceiling is proven, and reach for full fine-tuning only when the behavior you need cannot be reached with adapters. Fine-tuning is a tool, not a strategy; the strategy is a working production loop that improves every quarter.
Nate Nead is the CEO of Nead, LLC, a specialized AI development and private AI implementation firm. He and his team help organizations design, build and deploy private AI systems, including on-premise and private-cloud language models, retrieval pipelines and AI agents, that keep sensitive data under the client's control. Over more than 15 years, Nate has founded and grown more than a dozen technology and services brands, and has worked with teams at some of the most well-recognized companies in the Fortune 1000.
Bringing AI in-house, the right way.
Talk through your private or on-prem LLM deployment with an expert who has shipped them in regulated environments.
Private AI, in your inbox.
Occasional, high-signal notes on enterprise LLM deployment, security, and model strategy. No spam.














