Fine-Tuning vs RAG: How to Customize an AI Model
Introduction
Once a team decides to build a product around a large language model, one question comes up almost immediately: the general-purpose model is impressive, but it doesn't know our data, doesn't speak in our brand's voice, and doesn't follow our specific workflows — how do we actually customize it? Two techniques dominate the conversation: fine-tuning, which further trains the model's own internal parameters on a custom dataset, and Retrieval-Augmented Generation (RAG), which leaves the model's parameters untouched but feeds it relevant external information at the moment of answering a question.
These two approaches are frequently framed as competitors, and teams often ask which one they should choose. That framing, while understandable, is usually the wrong way to think about the decision. Fine-tuning and RAG solve genuinely different problems, and the right question isn't "which is better" but "what am I actually trying to change about the model's behavior." This article walks through what each technique actually does under the hood, where each one shines, where each one falls short, how to combine them, and — with concrete code — what building each one actually looks like in practice.
1. The Core Distinction: Changing Knowledge vs. Changing Behavior
The cleanest way to think about the difference between fine-tuning and RAG is this: RAG changes what information the model has access to. Fine-tuning changes how the model behaves, reasons, and communicates.
A pretrained language model already has an enormous amount of general knowledge and general capability baked into its parameters from training. When you want it to answer questions using facts it was never trained on — your company's product catalog, this morning's news, a customer's specific order history — you need to get that new information in front of the model somehow, and RAG is the standard way to do that: retrieve the relevant facts and hand them to the model as context.
When you want the model to fundamentally change its behavior — always respond in a very specific format, adopt a distinctive tone and voice consistently across every interaction, follow a complex, multi-step internal workflow specific to your business, or reliably perform a narrow, specialized task (like extracting structured data from messy invoices in a very particular way) far more consistently than generic prompting alone can achieve — fine-tuning is the tool built for that job, because it actually adjusts the model's internal weights so that the desired behavior becomes, in a sense, the model's new default, baked in rather than something that has to be re-explained through instructions every single time.
This distinction — facts versus behavior — is not a hard, perfectly clean line in every single case, and there is real overlap in what each technique can achieve to some degree. But it is the single most useful starting heuristic for deciding where to even begin, and confusing the two is the most common and costly mistake teams make when first approaching model customization: reaching for expensive, slow fine-tuning to solve a problem that a well-built RAG pipeline would have solved more cheaply and more flexibly, or conversely, trying to stuff enormous amounts of context into every prompt via retrieval to achieve a consistent behavioral change that fine-tuning would have handled far more reliably and cheaply per query.
2. How RAG Customizes a Model (Recap)
Retrieval-Augmented Generation, covered in much greater depth in the companion article dedicated specifically to it, works by leaving the underlying language model's weights completely unchanged and instead building a pipeline around it: relevant documents are split into chunks, converted into vectors using an embedding model, and stored in a searchable index. At query time, the user's question is embedded, the most relevant chunks are retrieved by similarity search, and those chunks are inserted directly into the model's prompt alongside the original question, so the model can generate an answer grounded in that retrieved material.
def rag_customize(query, document_index, embed_fn, generate_fn, top_k=4):
"""The model itself is untouched -- customization happens entirely
in what gets retrieved and stuffed into the prompt."""
query_vector = embed_fn(query)
relevant_chunks = retrieve_top_k(query_vector, document_index, k=top_k)
prompt = f"""Use the following context to answer the question.
Context:
{format_chunks(relevant_chunks)}
Question: {query}
Answer:"""
return generate_fn(prompt) # same base model, every single time
The key architectural property to notice here: the underlying model — whatever weights it was shipped with — never changes. Every ounce of customization lives entirely in the surrounding pipeline: what gets indexed, what gets retrieved, and how the prompt is constructed. This has significant, practical implications we'll return to throughout this article.
3. How Fine-Tuning Customizes a Model
Fine-tuning takes a different approach entirely: rather than leaving the model's weights fixed and working around them, it continues the training process itself, adjusting the model's internal parameters using a custom dataset of examples that demonstrate the desired behavior.
Preparing a fine-tuning dataset
The core input to fine-tuning is a dataset of example input-output pairs — typically formatted as example conversations or example prompt-and-ideal-response pairs — that demonstrate the behavior you want the model to adopt consistently. If you want a customer support model that always responds in a specific, calm, brand-appropriate tone and always follows a specific three-step troubleshooting structure, your fine-tuning dataset would consist of many examples of customer questions paired with exactly that kind of ideal, structured, appropriately-toned response.
# A simplified example of what a fine-tuning dataset entry might look like,
# in a common "chat-style" format used by many fine-tuning APIs.
fine_tuning_example = {
"messages": [
{"role": "system", "content": "You are a support agent for Acme Corp. "
"Always respond calmly, acknowledge the "
"customer's frustration, and end with a "
"clear numbered next step."},
{"role": "user", "content": "My order still hasn't arrived and it's been "
"two weeks!"},
{"role": "assistant", "content": "I completely understand how frustrating "
"that delay must be, and I'm sorry for "
"the wait. Here's what I'll do:\n"
"1. I'm checking the tracking status now.\n"
"2. If it's lost in transit, I'll issue "
"a replacement immediately.\n"
"3. I'll follow up with you within 24 hours."}
]
}
# A real fine-tuning dataset is typically hundreds to many thousands of
# examples like this one, saved as a JSONL file (one JSON object per line).
def write_finetune_dataset(examples, path):
import json
with open(path, "w") as f:
for example in examples:
f.write(json.dumps(example) + "\n")
The training process, conceptually
Once a dataset like this has been assembled, the fine-tuning process itself, at a conceptual level, works the same way the underlying model's original pretraining did (as described in the companion article on how large language models generate text): the model is shown each example, its actual generated response is compared against the ideal response in the dataset, and its weights are nudged, ever so slightly, via backpropagation and gradient descent, in the direction that would have made it more likely to produce that ideal response on its own. Repeated across the entire dataset, often for several passes (called epochs), this process gradually reshapes the model's behavior to more consistently match the patterns demonstrated in the fine-tuning examples.
In practice, teams rarely implement this raw training loop themselves from scratch when using a major model provider's fine-tuning service — most providers expose a simple API where you upload your formatted dataset and the underlying training infrastructure is handled for you.
# A representative example of what kicking off a fine-tuning job
# looks like via a typical provider API (illustrative, not tied to
# any one specific vendor's exact interface).
def start_finetuning_job(client, training_file_path, base_model, hyperparameters=None):
uploaded_file = client.files.upload(
file=training_file_path,
purpose="fine-tune"
)
job = client.fine_tuning.jobs.create(
training_file=uploaded_file.id,
model=base_model,
hyperparameters=hyperparameters or {"n_epochs": 3}
)
return job.id
# Once training finishes, the provider gives you a new model identifier
# representing your fine-tuned model, which you call exactly like the
# base model, except its behavior now reflects your training examples.
def use_finetuned_model(client, finetuned_model_id, user_message):
response = client.chat.completions.create(
model=finetuned_model_id,
messages=[{"role": "user", "content": user_message}]
)
return response.choices[0].message.content
Parameter-efficient fine-tuning
Fully retraining every single parameter of a large model is extremely computationally expensive and requires very large, carefully curated datasets to avoid problems like the model forgetting general capabilities it had before fine-tuning (a phenomenon called catastrophic forgetting). A widely used alternative, called parameter-efficient fine-tuning, and specifically a popular technique within that family called LoRA (Low-Rank Adaptation), instead freezes almost all of the original model's weights and only trains a small number of additional, newly-introduced parameters layered alongside the original ones. This dramatically reduces the computational cost and data requirements of fine-tuning, while still achieving much of the behavioral customization that full fine-tuning would provide, which is a major reason parameter-efficient techniques have become the dominant, practical way most organizations actually perform fine-tuning today, rather than the much more resource-intensive alternative of updating every single one of a model's original parameters directly.
4. Head-to-Head: Where Each Technique Actually Wins
When RAG is the clear right choice
The information changes frequently. If the underlying facts update daily, weekly, or even in real time — stock prices, inventory levels, breaking news, a constantly-updated support knowledge base — RAG handles this naturally, since updating the retrieval index is as simple as re-embedding the changed documents. Fine-tuning a model every time underlying facts change would be prohibitively slow and expensive, and there would always be a window where the model's baked-in "knowledge" from its last fine-tuning run has already gone stale.
You need traceability and citations. Because RAG explicitly retrieves specific source documents and can be prompted to cite them, it's straightforward to show a user exactly which document an answer was based on — valuable for trust, for compliance in regulated industries, and for letting users verify a claim themselves. Fine-tuned knowledge, baked into a model's weights, has no equivalent, inspectable paper trail; you cannot ask a fine-tuned model to show you the specific training example that led to a particular fact in its answer.
The information is private, sensitive, or access-controlled per user. RAG's retrieval step can be filtered per query based on what a specific user is authorized to see (as discussed in depth in the companion article on RAG), which is architecturally straightforward. Baking sensitive, user-specific information into a fine-tuned model's weights, by contrast, means that information becomes part of the model itself — available, in principle, to every single user who has access to that model, with no natural mechanism for restricting it per-request, which is a serious and often disqualifying limitation for handling genuinely sensitive or access-controlled data through fine-tuning alone.
You want to avoid the cost, complexity, and risk of retraining. Fine-tuning requires assembling a quality dataset, running and monitoring a training job, evaluating the resulting model to make sure it didn't degrade in unexpected ways, and potentially repeating this process whenever the desired behavior needs updating. RAG's iteration loop is comparatively lightweight: update the document index, tune the retrieval and prompt, and you're done, with no retraining cycle required at all.
When fine-tuning is the clear right choice
You need a highly specific, consistent output format or behavior, at scale, cheaply, on every request. If you need a model to reliably extract structured data from documents in an extremely particular schema, every single time, with high accuracy and without needing extensive formatting instructions repeated in every prompt, fine-tuning bakes that behavior in directly, producing more consistent results with shorter, cheaper prompts than trying to achieve the same reliability through lengthy in-context instructions alone.
You want the model to internalize a specific reasoning pattern or skill, not just have access to specific facts. Teaching a model to reliably reason through a specialized, multi-step domain-specific analysis — a particular style of legal contract review, for instance, or a specific diagnostic reasoning pattern used in a technical support workflow — is a behavioral, procedural change that fine-tuning is well suited to, since it's less about supplying new facts and more about shaping how the model approaches a category of problem.
Prompt length and cost are a serious constraint. If achieving the desired behavior via prompting alone would require an extremely long, detailed system prompt repeated on every single request (driving up both cost and latency for every query), fine-tuning can internalize much of that instruction directly into the model's weights, allowing for meaningfully shorter prompts at inference time while still achieving the same or better behavioral consistency.
You need a distinctive, consistent voice or style across an enormous volume of generated content, where achieving that consistency purely through prompting has proven unreliable or requires an unwieldy amount of in-context guidance to maintain across many different types of queries and edge cases.
Latency is critical and retrieval adds unacceptable overhead. RAG's retrieval step adds some latency to every single query — usually modest, but not zero. For latency-critical applications where even that added overhead genuinely matters, and where the underlying required behavior doesn't actually depend on fast-changing external facts, a fine-tuned model responding directly, without a retrieval step in the loop at all, can be meaningfully faster.
5. Combining Both: The Common Real-World Pattern
In practice, many of the most sophisticated production systems don't choose one technique exclusively — they combine both, using each for what it does best. A fine-tuned model handles the "how" — consistently adopting the right tone, format, and reasoning approach for the specific application — while a RAG pipeline handles the "what," feeding that fine-tuned model with current, specific, retrievable facts at query time.
def combined_pipeline(query, document_index, embed_fn, finetuned_generate_fn, user_groups):
"""Combine RAG's fresh, retrievable facts with a fine-tuned model's
consistent behavior, tone, and reasoning style."""
query_vector = embed_fn(query)
relevant_chunks = retrieve_with_access_control(
query_vector, document_index, allowed_groups=user_groups, top_k=4
)
prompt = f"""Context:
{format_chunks(relevant_chunks)}
Question: {query}
Answer:"""
# The model itself has already been fine-tuned to respond in the
# organization's specific tone, format, and reasoning style --
# RAG simply supplies it with the specific facts it needs for this query.
return finetuned_generate_fn(prompt)
A concrete example of why this combination works so well: imagine a customer support system for a software company. Fine-tuning teaches the model to consistently structure every response with an empathetic opening, a clear numbered troubleshooting sequence, and a specific sign-off — the behavioral pattern the company wants in every single interaction, regardless of topic. RAG then supplies the specific, current facts needed to fill in that structure correctly for any given question — the actual, up-to-date troubleshooting steps for a particular product feature, pulled from a knowledge base that engineers update independently of any model retraining cycle. Neither technique alone would achieve both goals as cleanly: RAG alone, on top of an unfine-tuned base model, would need very careful, lengthy prompting to reliably enforce the exact desired response structure on every single query, while fine-tuning alone, without RAG, would need to be retrained every time a troubleshooting step changed, which for a fast-moving software product could be uncomfortably often.
6. Cost, Time, and Maintenance Trade-offs
Beyond the functional differences described above, RAG and fine-tuning also differ substantially in their ongoing cost and maintenance profile, which is often just as important a factor in a real decision as which technique is functionally better suited to the problem.
Upfront cost and time. RAG typically has a lower upfront cost to get a first working version running — building a basic retrieval pipeline over an existing document collection can often be accomplished in days. Fine-tuning requires assembling a quality training dataset (frequently the most time-consuming part of the entire process, especially if suitable example data doesn't already exist and has to be created or curated from scratch), setting up and running a training job, and then rigorously evaluating the resulting model before deploying it — a process that typically takes meaningfully longer to get right the first time.
Ongoing maintenance. RAG's ongoing maintenance burden centers on keeping the document index synchronized with changing source material, which, once an automated pipeline for it exists, tends to be a comparatively lightweight, largely automatable ongoing cost. Fine-tuning's ongoing maintenance burden centers on periodically retraining as desired behavior needs to evolve, or as the underlying base model itself is upgraded to a newer version, which typically requires re-running the entire fine-tuning process against the new base model and re-validating the result, an inherently more manual and periodically recurring cost.
Per-query cost. A RAG pipeline adds the cost of an embedding call and a vector search to every query, on top of the underlying generation cost, and often increases the generation cost itself somewhat by adding retrieved context to the prompt. A fine-tuned model, once trained, typically has a per-query cost similar to or sometimes even lower than the equivalent base model (especially if fine-tuning allowed for shorter prompts by baking instructions into the weights instead), though the upfront training cost has to be amortized across enough queries to make that investment worthwhile.
7. A Practical Decision Framework
Given everything above, here is a practical way to approach the decision when starting a new customization project.
Start by asking: is the core problem that the model doesn't know something, or that it doesn't behave the way I want? If it's a knowledge gap — missing facts, outdated information, private data — start with RAG; it's cheaper, faster to iterate on, and naturally handles changing information. If it's a behavioral gap — inconsistent tone, unreliable format, an inability to reliably perform a specialized reasoning task even when given the relevant facts — consider fine-tuning, but first genuinely test whether better prompting or a longer, more carefully engineered system prompt closes the gap sufficiently, since prompting is far cheaper to iterate on than a full fine-tuning cycle and is frequently underestimated in how much behavioral consistency it can actually achieve on its own before fine-tuning becomes truly necessary.
If, after honestly assessing both, the problem genuinely spans both dimensions — you need both current, specific, retrievable facts and a highly consistent, specialized behavioral pattern — plan for the combined architecture described above from the outset, rather than building one piece in isolation and discovering the gap later. And regardless of which path you take, build in a real evaluation process before committing to either investment at scale: for RAG, this means testing retrieval quality and generation faithfulness against representative queries; for fine-tuning, this means evaluating the resulting model against a held-out test set of examples it wasn't trained on, specifically checking both for the desired behavioral improvement and for any unintended degradation of the model's general capabilities that fine-tuning can sometimes introduce if not done carefully.
8. Common Myths About Fine-Tuning and RAG
A number of misconceptions come up repeatedly enough in discussions of these two techniques that it's worth addressing them directly.
Myth: "Fine-tuning teaches the model new facts, so I can fine-tune it on my documents instead of building RAG." This is one of the most common and costly misunderstandings in the field. While fine-tuning technically can shift a model's outputs toward facts present in its training examples, it is a notoriously unreliable way to inject precise, retrievable factual knowledge. A model fine-tuned on a set of documents doesn't "memorize" them the way a database stores a record — it adjusts millions or billions of internal weights based on statistical patterns across the training examples, and there's no guarantee any specific fact will be reliably recalled correctly and completely later, especially for facts that appeared rarely in the fine-tuning data or that need to be reproduced precisely (an exact number, a specific date, a precise legal clause). RAG, by contrast, retrieves the actual source text verbatim and hands it directly to the model, which is a fundamentally more reliable way to ensure a specific fact is available and accurately represented at answer time.
Myth: "RAG is always cheaper than fine-tuning." This is true for upfront cost in many cases, but not universally true for total cost of ownership. A RAG system serving extremely high query volumes, with long retrieved contexts added to every single prompt, can accumulate substantial per-query costs over time that a fine-tuned model — potentially requiring much shorter prompts because relevant instructions and patterns are baked into its weights — might avoid. The right cost comparison has to account for expected query volume and prompt length, not just the simpler upfront-investment comparison.
Myth: "You should always start with fine-tuning because it's the 'more advanced' technique." Fine-tuning is not inherently more sophisticated or a natural "next step" up from prompting and RAG — it is a different tool suited to different problems, and reaching for it prematurely, before establishing that the actual bottleneck is behavioral rather than informational, is one of the most common ways teams waste significant engineering time and budget on a customization approach that doesn't address the actual root cause of a model's underperformance.
Myth: "Once you fine-tune a model, you're locked into that specific version forever." While it's true that fine-tuning has to be redone against a new base model when you want to upgrade to a newer underlying model version, this is a manageable, periodic maintenance cost rather than a permanent lock-in — most fine-tuning workflows, once a good dataset and evaluation process exist, can be re-run against an updated base model relatively efficiently, especially with parameter-efficient techniques like LoRA that require comparatively modest additional training time and computation per run.
Myth: "RAG makes fine-tuning entirely unnecessary now that context windows are so large." As discussed in the companion RAG-focused article, larger context windows reduce but don't eliminate the value of retrieval, and they address a completely different problem than fine-tuning does in the first place — a large context window lets you fit more retrieved facts into a single prompt, but it does nothing on its own to make a model consistently adopt a specific tone, format, or specialized reasoning pattern across every interaction, which remains squarely fine-tuning's domain regardless of how large context windows get.
9. A Worked Case Study: Building a Legal Document Assistant
To make the decision framework concrete, consider a realistic scenario: a legal team wants an AI assistant to help analyze incoming contracts, flag risky clauses, and answer questions about the firm's specific contract templates and internal review guidelines.
What actually needs RAG. The firm's specific contract templates change periodically as legal standards and internal policy evolve. The internal review guidelines document is updated by senior partners every few months. And any given user asking a question needs an answer grounded in the current version of these documents, with a clear citation back to the specific template clause or guideline paragraph the answer is based on, both for the legal team's own confidence in the tool and for audit purposes. This is a textbook RAG use case: facts that change over time, need traceability, and benefit from precise, verifiable retrieval rather than being baked into a model's weights where they'd inevitably grow stale and lose their citable source.
# Simplified sketch of the retrieval side of this system
legal_document_index = build_index(
documents=load_documents(["contract_templates/", "review_guidelines/"]),
embed_fn=legal_embedding_model # possibly a domain-fine-tuned embedding model
)
def answer_legal_question(query, user_role):
allowed_tags = get_access_tags_for_role(user_role)
chunks = hybrid_retrieve(query, legal_document_index, legal_embedding_model,
allowed_tags=allowed_tags, top_k=5)
prompt = build_cited_prompt(query, chunks)
return finetuned_legal_model.generate(prompt)
What actually needs fine-tuning. The firm wants every risk-flagging analysis to follow a very specific, consistent internal format — a structured breakdown by clause category, a standardized severity rating scale, and a particular tone that matches how the firm's own senior attorneys typically phrase risk assessments, distinct from how a generic model would phrase the same underlying analysis. This is a behavioral and stylistic consistency requirement, not a factual one, and is exactly the kind of thing fine-tuning handles well: training the model on a curated set of example contract excerpts paired with example risk analyses written in exactly the firm's preferred format and tone, until the model reliably reproduces that structure and style on new, unseen contracts, without needing that entire formatting specification re-explained via an enormous system prompt on every single request.
The combined system. Putting it together, the deployed assistant uses a model that has been fine-tuned specifically to produce risk analyses in the firm's exact preferred format and tone, and that fine-tuned model is fed, via RAG, the current, specific, correctly-cited contract template and guideline text relevant to whatever question or document is currently being analyzed. Neither piece alone would fully satisfy the requirement: RAG alone, without fine-tuning, would need constant, careful prompt engineering to enforce the firm's exact desired format on every query, with a real risk of inconsistency across different types of questions; fine-tuning alone, without RAG, would go stale every time internal guidelines were updated and would offer no reliable citation trail back to the specific source clause an analysis was based on — an unacceptable gap for a legal application where verifiability is not optional.
10. Frequently Asked Questions
Can I fine-tune a model and then also add RAG later? Yes, and this is a very common evolution path. Many teams start with a fine-tuned model addressing a behavioral requirement, and layer RAG on top later once they identify a separate need for current or private factual grounding, without needing to redo the fine-tuning work already completed — the two systems are architecturally independent and can be developed, and evolved, on separate timelines.
Does fine-tuning risk making a model worse at general tasks? Yes, this is a real risk called catastrophic forgetting, where a model fine-tuned heavily on a narrow dataset can lose some of its broader general capability or become overly rigid, producing the fine-tuned behavior even in situations where a different, more flexible response would have been more appropriate. This is mitigated by using a diverse enough fine-tuning dataset, using parameter-efficient techniques (which tend to cause less forgetting than full fine-tuning), keeping the number of training epochs modest, and evaluating the fine-tuned model against a broad test set, not just examples narrowly similar to the fine-tuning data, before deploying it.
How much data do I need to fine-tune effectively? This varies considerably depending on how narrow and well-defined the target behavior is and which fine-tuning technique is used, but useful, noticeable improvements have been achieved with datasets ranging from a few hundred high-quality, carefully curated examples up to many thousands, with parameter-efficient techniques like LoRA generally requiring less data than full fine-tuning to achieve a comparable degree of behavioral change. Quality and representativeness of the examples matters considerably more than sheer quantity — a smaller set of carefully chosen, genuinely representative examples covering the range of situations the model will actually encounter typically outperforms a much larger but noisier or less representative dataset.
Is there a middle ground between full fine-tuning and pure prompting? Yes — beyond the parameter-efficient fine-tuning techniques discussed above, teams often first exhaust the possibilities of careful prompt engineering (detailed system instructions, well-chosen few-shot examples embedded directly in the prompt) before concluding that fine-tuning is actually necessary, since a well-engineered prompt with good examples can often achieve much of the desired behavioral consistency at a fraction of the cost and complexity of a fine-tuning project, and is worth genuinely exhausting as an option first.
Can I use one model to fine-tune and a different model for embeddings in my RAG pipeline? Yes, and this is actually the norm rather than the exception — the embedding model used for retrieval and the generation model used to produce the final answer are typically entirely separate models, chosen independently based on what each is best suited for, and there is no requirement that they come from the same provider or share any architecture. In a combined fine-tuning-plus-RAG system, it's common to use an off-the-shelf, unmodified embedding model for retrieval while the generation model is the one that has been fine-tuned for behavior and tone.
What happens if I fine-tune on bad or biased examples? The resulting model will faithfully learn whatever patterns are present in the training data, including unwanted ones — biased phrasing, factual errors baked into example responses, or inconsistent formatting across examples will all be learned just as readily as the intended, desired patterns. This is exactly why the data curation and validation steps described earlier in this article are not optional best practices but a genuinely essential part of a responsible fine-tuning process, since a model has no independent way to distinguish a training example that represents the behavior you actually want from one that doesn't.
11. A Deeper Look: The Full Fine-Tuning Lifecycle
Because fine-tuning is a more involved, multi-stage process than standing up a basic RAG pipeline, it's worth walking through its full lifecycle in more detail than the earlier overview, since the quality of the outcome depends heavily on getting each stage right.
Defining success criteria before collecting data
Before writing a single training example, it's worth being explicit about what "success" actually looks like: what specific behavior should the fine-tuned model exhibit that the base model doesn't, and how will that behavior be measured afterward? Skipping this step is a common source of wasted effort — teams sometimes collect a large training dataset only to discover, after training, that they have no clear, objective way to determine whether the resulting model is actually better for their purposes than the original base model was, beyond a vague, subjective impression.
Collecting and curating training examples
Good training examples are harder to come by than teams often expect. They need to be genuinely representative of the range of inputs the model will see in production — not just the easy, typical cases, but also the edge cases, ambiguous situations, and less common variations that a model trained only on "clean" examples will handle poorly. Examples are often sourced from a combination of real historical data (past customer interactions, past documents processed by human experts) and deliberately authored synthetic examples designed to fill in gaps the real historical data doesn't adequately cover. Whatever the source, every example should be reviewed for quality — a training set contaminated with even a modest fraction of poor-quality, inconsistent, or simply incorrect examples can measurably degrade the resulting model's behavior, since the model has no way to distinguish a "good" example from a "bad" one in the data it's shown; it simply learns the patterns present in whatever it's given.
def validate_training_example(example):
"""A lightweight sanity-check pass before an example goes into the dataset."""
messages = example.get("messages", [])
if len(messages) < 2:
return False, "Example must have at least a user message and a response"
if not any(m["role"] == "assistant" for m in messages):
return False, "Example must include an assistant response"
assistant_msgs = [m["content"] for m in messages if m["role"] == "assistant"]
if any(len(msg.strip()) < 10 for msg in assistant_msgs):
return False, "Assistant response looks too short to be a useful example"
return True, "ok"
def curate_dataset(raw_examples):
valid, rejected = [], []
for example in raw_examples:
ok, reason = validate_training_example(example)
(valid if ok else rejected).append((example, reason))
return valid, rejected
Splitting data for honest evaluation
Just as with any machine learning workflow, it's essential to hold out a portion of curated examples — never shown to the model during training — specifically to evaluate the fine-tuned model afterward. Evaluating a model only against examples it was trained on gives a falsely optimistic picture of how well it will generalize to new, unseen inputs in actual production use, since the model may simply be reproducing memorized patterns from training rather than genuinely having learned the underlying, generalizable behavior.
import random
def split_dataset(examples, eval_fraction=0.15, seed=42):
shuffled = examples.copy()
random.Random(seed).shuffle(shuffled)
split_point = int(len(shuffled) * (1 - eval_fraction))
return shuffled[:split_point], shuffled[split_point:] # train_set, eval_set
Running training and monitoring for problems
During the actual training run, it's important to monitor for signs of overfitting — where the model's performance on the held-out evaluation set stops improving, or starts getting worse, even as its performance on the training examples themselves keeps improving, indicating the model has started memorizing quirks specific to the training set rather than learning genuinely generalizable behavior. Most fine-tuning platforms surface a training loss curve and, if an evaluation set is provided, a separate validation loss curve, and watching for validation loss beginning to rise while training loss continues falling is one of the clearest, most actionable signals that training should be stopped, or that the number of training epochs should be reduced in a subsequent run.
Post-training evaluation
Once training completes, the resulting model needs to be evaluated against the held-out evaluation set, ideally using the same success criteria defined at the very start of the process. This evaluation should check not just whether the desired new behavior is present, but also whether the model's more general capabilities have been preserved — running it against a broader, more general test set unrelated to the specific fine-tuning goal, to catch any unwanted degradation (the catastrophic forgetting problem discussed earlier) before the model is deployed to real users.
def evaluate_finetuned_model(model, eval_examples, judge_fn):
"""judge_fn scores a model's response against the ideal response,
e.g. using a rubric, exact-match check, or another model as a judge."""
scores = []
for example in eval_examples:
user_message = next(m["content"] for m in example["messages"] if m["role"] == "user")
ideal_response = next(m["content"] for m in example["messages"] if m["role"] == "assistant")
actual_response = model.generate(user_message)
score = judge_fn(actual_response, ideal_response)
scores.append(score)
return sum(scores) / len(scores) if scores else 0.0
12. A Deeper Look: Tuning a RAG Pipeline for Quality
Just as fine-tuning has its own multi-stage lifecycle, iterating on a RAG pipeline's quality involves its own set of concrete, testable levers, worth walking through with the same level of practical detail.
Building a retrieval test set
The single most valuable artifact for improving a RAG pipeline's quality is a curated set of representative test queries, each paired with the specific document or chunk known to actually contain the correct answer. This test set makes it possible to objectively measure retrieval quality — what fraction of test queries actually surface the correct chunk within the top-k results — rather than relying on a vague, subjective impression of "the answers seem pretty good" that doesn't reliably catch subtle regressions when a chunking strategy, embedding model, or retrieval parameter is changed.
def evaluate_retrieval(test_cases, index, embed_fn, top_k=4):
"""test_cases: list of (query, expected_doc_id) pairs."""
hits = 0
for query, expected_doc_id in test_cases:
retrieved = retrieve(query, index, embed_fn, top_k=top_k)
if any(chunk["doc_id"] == expected_doc_id for chunk in retrieved):
hits += 1
return hits / len(test_cases)
Iterating on chunk size systematically
Rather than guessing at a chunk size once, running the retrieval evaluation above across a handful of different chunk sizes and overlap settings makes it possible to empirically identify which configuration actually performs best for a specific document collection and query pattern, rather than relying on a generic rule of thumb that may not fit the specific characteristics of the content in question.
def sweep_chunk_sizes(documents, test_cases, embed_fn, chunk_sizes=(200, 400, 600, 800)):
results = {}
for size in chunk_sizes:
index = build_index(documents, embed_fn, chunk_size=size)
results[size] = evaluate_retrieval(test_cases, index, embed_fn)
return results # {chunk_size: retrieval_accuracy}
Comparing embedding models
Because different embedding models produce different, generally incompatible vector spaces, comparing them requires rebuilding the entire index for each candidate model and re-running the same retrieval evaluation, but this is exactly the kind of empirical comparison worth doing before committing to a specific embedding model in production, since performance can vary meaningfully between models, especially for specialized or non-English content.
13. Governance, Compliance, and Auditability Considerations
For organizations in regulated industries — healthcare, finance, legal services — the choice between fine-tuning and RAG carries governance implications beyond pure technical performance that are worth calling out explicitly.
RAG's inherent traceability — the ability to point to the specific source document and passage an answer was drawn from — tends to align well with regulatory and audit requirements that demand explainability: a compliance officer reviewing why an AI system gave a particular answer can be shown the actual retrieved source text, not just the model's own claim about where it got its information. This also makes it comparatively straightforward to update the system's behavior in response to a regulatory change: update the source document, re-index it, and the system's answers reflect the new guidance immediately, with a clear, auditable record of exactly when the underlying source material changed.
Fine-tuning, by contrast, bakes behavior into model weights in a way that is considerably harder to audit after the fact — there is no straightforward way to point to "this specific training example is why the model said that," and demonstrating to a regulator or auditor exactly why a fine-tuned model behaves the way it does on a specific input can require considerably more indirect evidence, such as the training dataset itself, evaluation results, and documented rationale, rather than a direct, inspectable causal chain. Organizations that do rely on fine-tuning in regulated contexts typically compensate for this with rigorous documentation of the training dataset's provenance and composition, thorough pre-deployment evaluation, and ongoing monitoring of production behavior, precisely because the direct auditability that RAG offers by nature isn't available in the same way.
This governance dimension is a meaningful, sometimes decisive factor in the fine-tuning versus RAG decision for regulated use cases, independent of which technique might otherwise seem like the better technical fit based purely on the knowledge-versus-behavior framework discussed earlier in this article.
14. Team and Organizational Considerations
Beyond the purely technical trade-offs, the two approaches also differ in what kind of ongoing organizational capability and process they demand, which is worth factoring into a realistic decision.
A RAG pipeline's ongoing health depends heavily on the quality and currency of the underlying document collection it draws from — which means it benefits enormously from a clear, established process (and, ideally, a specific owner) responsible for keeping that source content accurate, well-organized, and up to date. A RAG system built on top of a neglected, poorly organized, or rarely updated document collection will only ever be as good as that underlying collection allows, no matter how well the retrieval and generation machinery around it is engineered.
A fine-tuning program's ongoing health depends on having a repeatable process for collecting, curating, and validating new training examples as desired behavior evolves, along with genuine in-house (or vendor-supported) capability to run, evaluate, and monitor training jobs responsibly — including catching problems like catastrophic forgetting before a degraded model reaches production. Teams considering fine-tuning for the first time should honestly assess whether they have, or are willing to build, this kind of ongoing evaluation discipline, since a fine-tuning program undertaken without it tends to accumulate quiet, hard-to-detect quality regressions over successive training iterations, gradually eroding trust in the system in ways that can be difficult to trace back to their actual root cause once several rounds of retraining have occurred without a rigorous, consistent evaluation discipline in place from the start.
Summary
Fine-tuning and RAG are not competing solutions to the same problem — they customize different aspects of a language model's behavior, and the clearest way to choose between them is to ask what you're actually trying to change: RAG supplies a model with specific, current, or private facts at the moment of answering, by retrieving relevant content and inserting it into the prompt, leaving the model's own weights untouched; fine-tuning actually adjusts those weights through further training on a custom dataset, reshaping how the model behaves, reasons, and communicates by default, without needing that behavior re-specified through lengthy instructions on every single query. RAG tends to win when information changes frequently, needs to be traceable to its source, or must be access-controlled per user, and it's typically cheaper and faster to get started with and to maintain. Fine-tuning tends to win when the goal is a highly consistent behavioral pattern, a specialized reasoning skill, or a shorter, cheaper prompt at inference time, at the cost of a more involved upfront investment in dataset creation, training, and evaluation. And in many of the most capable real-world systems, the two are combined deliberately: a fine-tuned model providing consistent behavior and tone, fed by a RAG pipeline supplying the current, specific facts that behavior needs to be applied to — each technique doing the part of the job it's actually best suited for. When in doubt, start with the cheaper, faster-to-iterate option — usually RAG, or simply better prompting — and reach for fine-tuning only once you can point to a specific, well-defined behavioral gap that prompting genuinely cannot close, backed by an evaluation process rigorous enough to confirm the fine-tuned result is actually better, not just different. That discipline — grounding the decision in a concrete gap and a real before-and-after measurement, rather than intuition or the appeal of a more technically involved solution — is what separates a customization effort that reliably improves a product from one that simply adds cost and complexity without a correspondingly clear payoff — and it applies just as much to choosing between fine-tuning and RAG as it does to any other significant engineering investment a team might consider building on top of a large language model, whether that's the very first customization decision a project makes or the tenth iteration of an already-deployed system.

Comments
Post a Comment