rumblr Work in progressWIP

● The AI Primer · Lesson 11 · Part 1: how the model works inside

Fine-tuning in practice

preparing data, forgetting old skills, merging models

You'll be able to explain Preparing data, forgetting old skills, merging models

Members · open during launch 54 min22 figures and diagrams10 interactive
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Decide with an eval: build a held-out set first; try prompting, then retrieval, and fine-tune only for behaviour a prompt can't pin down. It pays off after C / (c_prompt − c_tuned) requests.
  2. Data: use the model's chat template and train only on assistant turns; deduplicate (Jaccard on word shingles); freeze the held-out set before training and remove near-copies of it; audit labels, because wrong labels cap what the eval can show. Quality beats quantity.
  3. Forgetting: every weight is shared, so learning a new task erodes old ones, in proportion to how far the weights move and how much the tasks conflict. Fewer steps, a lower learning rate and LoRA shorten the trip; replaying a little old data is what keeps a contradicted skill.
  4. Overfitting: on small data, validation loss bottoms out early; ship the best checkpoint, not the last.
  5. Merging: a task vector is θ_ft − θ_base. Adding task vectors combines skills without training; averaging is λ = 1/T and dilutes them; vectors that point in opposite directions interfere.

Level 1

The practitioner's guide

Start from what is on the table: the map picks the option, and the guide below explains every one.

In one sentence

Fine-tuning in practice is the work around the training run: deciding with an evaluation set whether to train at all, preparing examples that teach what you mean, keeping the skills the model already had, stopping before it memorises your data, and sometimes combining two fine-tunes by arithmetic on their weights instead of training a third.

When you need it

You need this lesson when a fine-tune is on the table: a behaviour the best prompt cannot make consistent, or a long prompt sent so often that its tokens are most of the bill. The tell for the first is an eval score that stops improving however the prompt is reworded; the tell for the second is arithmetic. In this lesson's worked example, a 3,000-token prompt at $2 per million tokens costs $0.0060 per request, a tuned model that needs 300 tokens at $4 per million costs $0.0012, and a $600 fine-tune pays for itself after 125,000 requests (25 days at 5,000 a day; the prices are illustrative). You don't need fine-tuning to add facts, which go stale the day they change and belong in retrieval, and you don't need it while a clearer instruction with two examples still moves the eval. Nothing in this lesson matters until the eval exists: it is the first box of the decision, before any model.

Your options

From the cheapest to the most committed:

Option What it does What it guarantees What it costs Where it lives
The best prompt, scored on the eval Instructions and examples in the prompt, measured on a frozen held-out set A baseline every other option must beat; survives a base-model upgrade Tokens on every call Your prompt
Hosted supervised fine-tuning Upload chat examples; the vendor applies the chat template, masks the user turns, trains and hosts Consistent format and style without the long prompt Curated examples, a training job, often a higher per-token price The vendor's fine-tuning API
LoRA on an open model Train a small adapter beside frozen weights on your examples Learns the task with less forgetting of the base model's skills (Biderman et al., 2024) A GPU, data, a serving stack; may learn a hard new domain less completely Your training stack
Full fine-tune, with replay Update every weight on your task mixed with a little general data The largest change the model can make, and the old skills kept if you replay them Multi-GPU training, the most forgetting when you don't replay, a copy of the model per variant Your training stack
Merge existing fine-tunes Add each fine-tune's task vector (its change from the base) to the base, no training Two separable skills in one model at no extra inference cost An eval on every task; fails when the two changes fight Your weights, with a merging toolkit

How to choose

Climb down only when the eval proves the rung above falls short.

  • A format or a house style: a few dozen to a few hundred excellent examples, hosted or LoRA. Start with 50 to 100 and double only while a doubling still beats the eval's margin of error (section 2e).
  • A behaviour that contradicts the base model ("always JSON" against "chat naturally", "be terse" against "explain"): expect the old behaviour to vanish wherever your data overrides it, and mix in general instruction-following data so the model stays an assistant (section 3c).
  • A harder specialised skill: thousands of examples, and consider a full fine-tune with replay, since LoRA learns less on a demanding domain.
  • Two skills trained separately, by different teams or on data that cannot be pooled: try a merge first, and check the cosine between the task vectors before you trust it (section 5b).
  • Whatever you pick, the fine-tune ships only if it beats the prompt on the same held-out set by more than that set's margin of error, and a general eval sits beside the task eval.

What it costs

Money is the one-off cost of writing and checking data plus the run, against a per-request saving that may be zero if the tuned model is priced like the base one; the break-even rule in section 1 puts a number on it. Every new base model repeats the one-off cost. Data is the expensive part, and the eval is dearer than people expect: 100 held-out examples measured at 80% mean "somewhere between 72% and 88%", and it takes 400 to halve that margin to ±3.9 points (this lesson's margin_of_error). Labels cost accuracy too: with 10% of reference labels wrong, a perfect model scores 90% and a 90% model scores 82%. Training itself is short. In the lesson's toy the new task is learned in 5 steps and every step after only erodes the general skill, and fine-tunes on small data run for a few epochs, with the best checkpoint shipped rather than the last.

What breaks

  • The wrong chat template. Training succeeds, the loss falls, and the deployed model sees role markers it never learned. Use the exact template the base model was trained with, and the system prompt you will deploy.
  • Training on the user's turns. Forget the loss mask and the model learns to write questions. Hosted APIs mask for you; render_chat shows which segments count.
  • A leaking eval. A training example that nearly copies a held-out one turns the eval into a memory test. Freeze the held-out set first and remove near-copies from training (Jaccard on word pairs, threshold 0.7 in the toy).
  • Forgetting. Fine-tuning on one task takes the toy's general skill from 0.99 to 0.655, and a task that contradicts an earlier one takes the earlier one from 0.98 to 0.025. A lower learning rate only slows the slide; 10 replayed examples in 210 bring it back to 0.945.
  • Memorising a small dataset. On 16 examples with 3 wrong labels, validation loss bottoms out at epoch 70 (0.319) and has quadrupled (1.277) by epoch 1,500. Checkpoint every epoch and ship the best.
  • A merge that cancels. Task vectors with a clearly negative cosine (−0.25 in the toy) leave at least one skill at a coin flip whatever the scale. Retrain jointly, or use a conflict-resolving merge.

In the wild

Hosted fine-tuning APIs take chat-formatted examples and handle the template and the mask; OpenAI's model optimization guide lists supervised fine-tuning, DPO and reinforcement fine-tuning. For open models, Hugging Face TRL's SFTTrainer runs the supervised loop and PEFT supplies LoRA. Merging has its own toolkit: mergekit implements linear averaging, SLERP, task arithmetic, TIES, DARE and more. The evidence behind the advice is in the papers at the end of the lesson: LIMA's 1,000 curated examples, model soups for averaging fine-tunes of one task, task arithmetic for adding and subtracting them, TIES for resolving interference, elastic weight consolidation for when the old data is gone, and Biderman et al. on LoRA learning less and forgetting less.

Go deeper

Level 2 puts every claim above on a 97-weight model you can train in a fraction of a second: the break-even formula, the chat mask, Jaccard deduplication, the margin of error and the label-noise ceiling, forgetting measured as distance from the base, the overfitting curve with early stopping, and task vectors merged at every scale. If you only needed to decide whether and how to fine-tune, you are done.

Level 2

How it works, from scratch

Picture a skilled cook who joins your restaurant. They already know how to cook (that is the pretrained model). You want them to cook your menu, your way. You can hand them a note with every order (a prompt), give them the recipe binder to look things up in (retrieval, RAG), or send them on a course about your kitchen (fine-tuning). The course is the only option that changes the cook, and it comes with the risks every teacher knows: the lessons can be badly written, the cook can cram the practice exam instead of learning, and a month of drilling one cuisine can make them rusty at everything else.

Everything below is one of those risks, measured on a model small enough to train in a fraction of a second.

Chapter 1

Should you fine-tune at all?

Put your own numbers in first; the worked example below will then read as what it is, one row of this calculator.

Everyday picture Giving the cook a note with every order costs a little each time. The course costs a lot once, and it has to be repeated whenever you hire a new cook (switch to a newer base model). The course pays off only when the notes would otherwise be long, sent millions of times, or simply not enough to make the cook consistent.

Tiny worked example A support bot sends a 3,000-token prompt full of instructions and examples on every request. A fine-tuned model has learned that behaviour and needs only 300 tokens of prompt. Take these illustrative prices (real ones vary by provider and change often): $2 per million input tokens for the general model, $4 per million for the tuned one, because hosting a custom model usually costs more per token.

Tokens per request Price per million Cost per request
prompted 3,000 $2 3,000 × 2 / 1,000,000 = $0.0060
fine-tuned 300 $4 300 × 4 / 1,000,000 = $0.0012

Each request saves $0.0048. Writing and checking 1,000 training examples plus the training run costs, say, $600 once. After 600 / 0.0048 = 125,000 requests the course has paid for itself: 25 days at 5,000 requests a day.

Figure 2 · Diagram

Reading it: start at the top, and notice that the first box is not a model at all: it is the evaluation set, the fixed exam every option is scored on. Each rung is cheaper and faster to change than the one below it, so you climb down only when the eval proves the rung above falls short. Missing knowledge goes to retrieval, because fine-tuning is a poor way to add facts; wrong behaviour that no prompt can pin down goes to fine-tuning. The last diamond matters most: a fine-tune ships only if it beats the prompt on the same eval. primer.ml.training_stages has the same decision as a flowchart of approaches; this one adds the measurement at every step.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
"N-star": the break-even number of requests 125,000
one-off cost: writing and checking data, the training run $600
cost of one request
cost of one request with the long prompt, and with the fine-tuned model $0.0060, $0.0012
input tokens sent per request 3,000 and 300
price in dollars per million input tokens $2 and $4
one million: prices are quoted per million tokens

In words: "divide what the course costs once by what it saves on each request; that is how many requests it takes to earn the cost back."

With the numbers: = 3,000 × 2 / 10⁶ = 0.006, = 300 × 4 / 10⁶ = 0.0012, so = 600 / 0.0048 = 125,000.

Level 3: in Python
# c = t·p / 1,000,000 for each option
c_prompt = 3000 * 2.0 / 1_000_000
c_tuned = 300 * 4.0 / 1_000_000
c_prompt, c_tuned  # → (0.006, 0.0012)
C_once = 600
# N* = C_once / (c_prompt − c_tuned)
N_star = C_once / (c_prompt - c_tuned)
round(N_star)  # → 125000
# days to break even at 5,000 requests a day
round(N_star / 5000, 1)  # → 25.0

Figure 1 · Drawn from the lesson's code

0 50000 100000 150000 200000 250000 requests served 0 200 400 600 800 1000 1200 1400 total cost so far ($) When does a fine-tune pay for itself? break-even 125,000 requests long prompt: $0.0060 per request fine-tuned: $600 once + $0.0012 per request

The prompted line starts at zero and climbs steeply; the fine-tuned line starts at 600 dollars and climbs slowly; they cross at 125,000 requests

Reading it: the x-axis counts requests served and the y-axis is the total money spent so far. The prompted line starts at zero but climbs steeply, because every request pays for 3,000 tokens. The fine-tuned line starts at $600 (the one-off cost) and climbs slowly. Left of the dashed line, prompting is cheaper; right of it, the fine-tune is. If the tuned model cost as much per request as the prompt, the two lines would never cross, and the only reason left to fine-tune would be quality.

In code: per_request_cost prices one request and break_even_requests returns , or infinity when the tuned model saves nothing per request.

Why it matters in practice. Three costs hide outside this formula. Every new base model means repeating the fine-tune, so recurs. Facts learned by fine-tuning go stale and are hard to update, which is why knowledge belongs in retrieval (primer.agents.rag). And the engineering time to build an eval is spent whichever way you go, which is why it comes first.

Chapter 2

Preparing the data

Edit two questions yourself first; section 2b then names what you counted.

Everyday picture Before the course, someone writes the course book. Every recipe must be written in the same layout, no recipe may appear five times, some recipes are locked in a drawer for the final exam before the cook ever sees the book, and the recipes must actually be right. Most fine-tuning failures are course-book failures.

Figure 4 · Diagram

Reading it: raw examples enter on the left and pass through five filters. The branch at "Set aside the held-out set" is the important one: the held-out examples leave the pipeline before anything is trained or tuned, and the step after it removes any training example that nearly copies one of them. The subsections below build each box.

2a. Formatting chat examples

Everyday picture A play script: every line starts with who speaks it. The model learns the play by reading scripts, and it is graded only on the lines of the character it will play, the assistant.

Tiny worked example One example as a list of role-tagged messages:

Turn Role Content Trained on?
1 system Be brief. no
2 user Capital of France? no
3 assistant Paris. yes

Rendered into text, it becomes <|system|>Be brief.<|end|><|user|>Capital of France?<|end|><|assistant|>Paris.<|end|>, and only the last segment counts towards the loss.

Figure 5 · Diagram

Reading it: the solid arrows are the order the model reads the turns. The dotted arrows carry no training signal: the system and user turns are context the model conditions on. Only the thick arrow feeds the loss, so the model learns to answer, not to write questions. This is the loss mask from primer.ml.training_stages, section 2.

The rules that matter. Use the exact chat template the base model was trained with; the role markers above are illustrative, and each model family has its own. Train with the same system prompt you will deploy with. Reject examples whose last turn is not the assistant's (nothing to learn), whose turns are empty, or whose roles are unknown.

In code: chat_example builds the message list, render_chat flattens it into (segment, trained?) pairs, and validate_chat lists every problem with an example.

Why it matters in practice. A template mismatch is silent: training succeeds, the loss falls, and the deployed model sees markers it never learned, so it behaves like the base model or worse.

2b. Deduplication

Everyday picture A flashcard deck with the same card in it five times. You study that card five times as often, and you start answering every question with it.

Tiny worked example Compare questions by their word pairs (two neighbouring words; a window of k words is called a shingle), after lowercasing and dropping punctuation:

Text Word pairs
"How do I reset my password?" how do, do i, i reset, reset my, my password
"how do I reset my password, please" the same 5, plus password please
"How do I change my email?" how do, do i, i change, change my, my email

The first two share 5 of the 6 distinct pairs between them: a near-duplicate. The first and third share 2 of 8: different questions.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the sets of word pairs of two texts 5 pairs and 6 pairs
intersection: pairs in both sets 5 shared pairs
union: pairs in either set, each counted once 6 distinct pairs
the number of items in a set
Jaccard similarity, from 0 (nothing shared) to 1 (identical) 0.83

In words: "count the pairs the two texts share, and divide by the number of distinct pairs they have between them."

With the numbers: J = 5 / 6 = 0.83 for the two password questions and 2 / 8 = 0.25 for password against email. A threshold of 0.7 keeps the first password question, drops its rewording, and keeps the email question.

Level 3: in Python
def pairs(text):
    words = text.lower().replace("?", "").replace(",", "").split()
    return {(a, b) for a, b in zip(words, words[1:])}
A = pairs("How do I reset my password?")
B = pairs("how do I reset my password, please")
C = pairs("How do I change my email?")
# |A ∩ B| and |A ∪ B|
len(A & B), len(A | B)  # → (5, 6)
# J(A, B)
round(len(A & B) / len(A | B), 2)  # → 0.83
# J(A, C)
len(A & C) / len(A | C)  # → 0.25

Figure 6 · Diagram

Reading it: examples arrive one at a time and are compared against everything already kept. A match above the threshold is dropped; anything else joins the kept list. The first copy of each group survives, so the order of the data decides which wording you keep.

In code: normalize lowercases and strips punctuation, shingles collects the word pairs, jaccard scores two texts and deduplicate returns the indices worth keeping. Comparing every pair is fine for thousands of examples; at web scale, MinHash estimates the same Jaccard without comparing every pair.

Why it matters in practice. Duplicates overweight a few examples, so the model parrots them. Worse, a near-copy of an eval question in the training set lets the model recite the answer, and the eval score becomes a memory test. Lee et al. found training sets with thousands of near-duplicates, and removing them made models memorise less.

2c. The held-out set comes first

Everyday picture A good teacher writes the final exam before teaching the course and locks it in a drawer. If the exam were written afterwards, it would drift towards what the class happened to practise.

Tiny worked example Fifty support questions are deduplicated, shuffled with a fixed seed, and 10 (20%) go in the drawer. Only then are the other 40 used, and any of the 40 that nearly copies one of the 10 is removed. Every later choice (prompt wording, learning rate, which checkpoint to ship) is scored on those 10. But how much can 10, or even 100, questions tell you?

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the accuracy measured on the held-out set 0.8
the number of held-out examples 100
the standard error: how much the measured accuracy would wobble if you drew a different held-out set of the same size 0.04
how many standard errors to allow; 1.96 covers 95% of the wobble 1.96
margin the true accuracy is probably within ± this of the measured one 0.078

In words: "the measured accuracy is uncertain by about two standard errors, and the standard error shrinks with the square root of the number of examples."

With the numbers: with 100 examples at 80%, the margin is 1.96 × √(0.8 × 0.2 / 100) = 1.96 × 0.04 = 0.078, so "80%" means "somewhere from about 72% to 88%". A fine-tune that scores 83% has not been shown to beat a prompt that scores 80%. With 400 examples the margin halves to 0.039.

Level 3: in Python
import math
a, n, z = 0.8, 100, 1.96
# the standard error sqrt(a(1 − a)/n)
round(math.sqrt(a * (1 - a) / n), 4)  # → 0.04
# the 95% margin
round(z * math.sqrt(a * (1 - a) / n), 4)  # → 0.0784
# four times as many examples halves it
round(z * math.sqrt(a * (1 - a) / 400), 4)  # → 0.0392

Figure 7 · Diagram

Reading it: the fixed seed makes the split repeatable, so everyone on a team scores against the same held-out set. The held-out branch is frozen the moment it is made. Both branches then meet in the leak check, which only ever removes examples from the training side.

Figure 3 · Drawn from the lesson's code

1 0 2 1 0 3 held-out examples n (log scale) 2 4 6 8 10 12 14 16 95% margin (percentage points) How precise is an 80% score? Four times the data halves the margin n = 100: ±7.8 n = 400: ±3.9 n = 1600: ±2.0

The margin of error falls from about 16 points at 25 examples to 1.4 points at 3,200; each fourfold increase halves it

Reading it: the x-axis is the number of held-out examples (log scale) and the y-axis is the ± margin, in percentage points, around a measured 80%. The marked points show the square-root law: 100 examples give ±7.8 points, 400 give ±3.9 and 1,600 give ±2.0. Each halving of the margin costs four times the examples, which is why held-out sets of a few hundred are common and a few dozen can only detect big differences.

In code: split_before_training deduplicates, shuffles with a seed, freezes the held-out set and removes leaks with remove_leaks; margin_of_error gives the ± for any accuracy and size.

Why it matters in practice. Every time you look at held-out scores and change something, the held-out set leaks a little into your decisions. A held-out set built first and used sparingly is the only honest measure of whether the fine-tune helped. See primer.ml.regularization for train, validation and test splits, and primer.agents.evals for building evals.

2d. Label quality

Everyday picture An answer key with typos. A student who gets every question right is marked wrong wherever the key is wrong, and a student who makes the same mistake as the key is marked right.

Tiny worked example A model is truly right 90% of the time, and 10% of the reference labels are wrong. It scores a point when it is right on a right label (0.9 × 0.9 = 0.81), or when it is wrong on a wrong label and the two mistakes cancel (0.1 × 0.1 = 0.01). The eval reports 82%, not 90%.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the model's true accuracy 0.9
epsilon: the share of reference labels that are wrong 0.1
right answer, right label 0.81
wrong answer on a wrong label (yes/no labels only, so the two errors agree) 0.01
what the eval reports 0.82

In words: "the eval gives credit for being right on a correct label, and by accident for being wrong on a wrong one."

With the numbers: 0.9 × 0.9 + 0.1 × 0.1 = 0.82. A perfect model () scores only 1 − ε = 0.9: noisy labels put a ceiling on what the eval can show.

Level 3: in Python
a, eps = 0.9, 0.1
# right on a right label, plus wrong on a wrong label
round(a * (1 - eps) + (1 - a) * eps, 2)  # → 0.82
# even a perfect model scores only 1 − ε
round(1.0 * (1 - eps), 2)  # → 0.9
# two labelers, five items: how often do they agree?
labeler_1 = ["yes", "no", "yes", "yes", "no"]
labeler_2 = ["yes", "no", "no", "yes", "no"]
sum(p == q for p, q in zip(labeler_1, labeler_2)) / len(labeler_1)  # → 0.8

Figure 8 · Diagram

Reading it: each item takes two coin flips: is the model right, and is the label right? Multiply along a path to get its share. Two paths are scored as right, 0.81 and 0.01, which add up to the 0.82 the eval reports. The 0.09 on the second path is a correct answer marked wrong by a bad label.

How to check label quality. Have two people label the same sample independently and measure how often they agree. Above, they agree on 4 of 5 items, 80%: one item in five is ambiguous or mislabelled, and your eval can't resolve differences smaller than that noise. Read every disagreement; they are usually unclear instructions, not careless labelers.

In code: measured_accuracy applies the formula and label_agreement compares two labelers.

Why it matters in practice. Bad training labels teach the model the mistakes (section 4 shows a small model memorising three of them). Bad eval labels hide real improvements. Zhou et al. (LIMA) fine-tuned a large model on just 1,000 carefully chosen examples and got a strong assistant: quality of examples beats quantity.

2e. How many examples?

Everyday picture Teaching a house style is not teaching a language. A new cook learns how you plate a dish from a few dozen good examples; they don't need ten thousand.

Tiny worked example Fine-tuning teaches a behaviour the base model can almost do already, so the useful sizes are small: a few dozen examples to show a format, hundreds for a consistent style or a narrow task, thousands for a harder specialised skill. Section 4 fine-tunes on 16 examples and shows the other side: with so few, the model soon memorises them, noise included.

Figure 9 · Diagram

Reading it: the loop grows the dataset only while each doubling still buys a real improvement, meaning one larger than the held-out margin from section 2c. When a doubling stops paying, more of the same data won't help; the fix is better examples or a different approach.

Why it matters in practice. Labelled data is the expensive part of fine-tuning. Doubling until the gains flatten spends it where it helps, and the held-out set tells you when to stop.

The margin from 2c, with the size of the held-out set in your hands.

The answer key from 2d, with its share of typos in your hands.

Chapter 3

Catastrophic forgetting

Train the chapter's toy model yourself first, one step at a time; the chapter then explains what you watched slide.

Everyday picture Someone who learned to drive in Britain, keeping left, moves to the United States and practises keeping right every day for a month. The new habit wins, and on a trip home they drift to the wrong side of the road. Nobody told them to forget the old rule; the new practice simply rewrote the reflex both rules use. Neural networks do this to an extreme: train on a new task alone and an old one can vanish. This is catastrophic forgetting.

The model we'll fine-tune. To watch it happen, we need a model small enough to train instantly. Each example is four numbers, and the answer is yes or no:

Task Inputs it uses Rule Plays the role of
general all four, anywhere in [−3, 3] yes when x₁ + x₃ > 0 the base model's pretraining
A x₁ in [−3, −1], x₂ in [−2, 2] yes when x₂ > 0 keep left in Britain
B x₁ in [1, 3], x₂ in [−2, 2] yes when x₂ < 0 keep right in the US
C x₃, x₄ in [−2, 2] yes when x₃ + x₄ > 0 an unrelated skill

A and B read the same two inputs and apply opposite rules in different regions, so one model can learn both, but only by paying attention to the region. C reads inputs that A and B never touch.

Figure 13 · Diagram

Reading it: four numbers go in, 16 hidden units each mix all of them, and one output unit turns the hidden units into a probability. Every weight (64 in the first layer, 16 hidden biases, 16 output weights and one output bias) sits in a single vector θ of 97 numbers. Every task flows through the same hidden units, which is exactly why one task's training can damage another. primer.ml.neural_net builds this kind of network from scratch.

The base model is this network trained on the general task (99% accuracy on held-out examples). Every fine-tune below starts from it and runs plain gradient descent on 200 examples.

In code: TinyNet holds θ and computes predictions, loss and TinyNet.gradient; make_task draws examples for each of the TASKS; base_model pretrains the base and fine_tune trains a copy, leaving the starting model untouched.

3a. The general skill fades while you teach a new one

Everyday picture A month of drilling one cuisine makes the cook a little rusty at everything else, and a second month of drilling it makes them rustier still, even though they had mastered the cuisine in the first week.

Tiny worked example Fine-tune the base on task A (learning rate 0.5) and check both skills on held-out examples:

After step Accuracy on A General skill Distance from base
0 (the base) 0.53 0.99 0
5 0.975 0.825 2.60
10 0.995 0.79 2.83
300 0.98 0.655 5.19

Task A is learned in 5 steps. The next 295 steps teach nothing new about A but keep eroding the general skill, from 0.825 to 0.655. The last column explains why: the weights keep travelling away from the base.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
theta: every weight of the base model, as one vector 97 numbers
every weight after fine-tuning steps
the number of weights 97 (3 in the hand example)
a counter over the weights 1 … P
the -th weight after steps
the length of vector : square each entry, add, take the square root
how far fine-tuning has moved the model 5.19 after 300 steps

In words: "the distance travelled is the length of the change in the weights: square every weight's change, add them up, take the square root."

With the numbers: for a three-weight model moving from (1, 0, 2) to (1.5, −1, 2), the changes are (0.5, −1, 0), and d = √(0.25 + 1 + 0) = √1.25 = 1.118.

Level 3: in Python
import math
theta_base = [1.0, 0.0, 2.0]
theta_t = [1.5, -1.0, 2.0]
# each weight's change
[t - b for t, b in zip(theta_t, theta_base)]  # → [0.5, -1.0, 0.0]
# ‖θ_t − θ_base‖
round(math.sqrt(sum((t - b) ** 2 for t, b in zip(theta_t, theta_base))), 3)  # → 1.118

Figure 14 · Diagram

Reading it: the gradient from task A's examples moves the shared weights, and both skills read those same weights. The dotted arrow is what is missing: the general skill has no examples in this training run, so nothing pushes back when a change that helps A hurts it. Forgetting is not an event; it is the absence of a counterweight.

Figure 10 · Drawn from the lesson's code

1 0 0 1 0 1 1 0 2 fine-tuning step (log scale) 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 held-out accuracy Learning A, forgetting the general skill task A, lr 0.5 general skill, lr 0.5 task A, lr 0.02 general skill, lr 0.02 0 1 2 3 4 5 distance from the base ‖θ − θ_base‖ 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 general-skill accuracy Forgetting tracks how far the weights move lr 0.5 lr 0.02

Left: accuracy on A jumps to 1 within a few steps while the general skill slides from 0.99 to 0.66 at learning rate 0.5 and only to 0.79 at 0.02. Right: the general skill falls as distance from the base grows

Reading it: on the left, the x-axis is the training step (log scale). Blue lines are accuracy on task A, red lines the general skill; solid is learning rate 0.5, dashed is 0.02. With the big rate, A is learned within a few steps and the general skill then slides for the rest of the run. With the small rate, A is learned later (around step 70) and the general skill ends at 0.79 instead of 0.655, with A just as good. On the right, every step of both runs is plotted as distance from the base against the general skill: the points fall along one downward curve. Forgetting tracks how far the weights move.

Mitigations that shorten the trip. Three knobs limit the distance, and the toy measures two of them:

  • Fewer steps. Stop once the held-out score on the new task stops improving. Here, stopping at step 10 keeps the general skill at 0.79 instead of 0.655, with A at 0.995.
  • A lower learning rate. Smaller steps travel less far for the same result: 0.79 instead of 0.655 after 300 steps.
  • A smaller update. LoRA (primer.ml.training_stages, section 4) freezes the base weights and allows only a low-rank change. Biderman et al. measured this on real language models and summed it up in their title: LoRA learns less and forgets less.

In code: general_skill_run fine-tunes on A and records, after every step, accuracy on A, the general skill and the distance from the base.

Why it matters in practice. A fine-tuned assistant that has become worse at everything outside its narrow task is the most common fine-tuning disappointment. Always score the general skills you care about alongside the new task, and prefer the earliest checkpoint that has learned the task.

3b. A new task that contradicts an old one

Everyday picture Back to the driver. Practising "keep right" does not merely add a skill; it pushes directly against "keep left", because both use the same reflex.

Tiny worked example Take the model fine-tuned on A (98% on A), then fine-tune it on B alone for 300 steps:

Second fine-tune A before A after New task after
on B (contradicts A) 0.98 0.025 1.00
on C (separate inputs) 0.98 0.97 0.98

After B, the model does not merely forget A; it answers A's questions backwards, because it learned "yes when x₂ < 0" everywhere. After C, A is barely touched: C's inputs never flowed through the weights A relies on most.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
held-out accuracy on the old task before the new fine-tune 0.98
the same, after the new fine-tune 0.025
forgetting: accuracy lost on the old task 0.955

In words: "forgetting is how much accuracy the old task lost."

With the numbers: F = 0.98 − 0.025 = 0.955 after B, and 0.98 − 0.97 = 0.01 after C.

Level 3: in Python
acc_before, acc_after_B, acc_after_C = 0.98, 0.025, 0.97
# F after the contradicting task B
round(acc_before - acc_after_B, 3)  # → 0.955
# F after task C, on separate inputs
round(acc_before - acc_after_C, 3)  # → 0.01

Figure 15 · Diagram

Reading it: both second fine-tunes start from the same model. The only difference is which weights the new task needs: B needs the very ones A uses, in the opposite direction; C mostly needs others. How much a model forgets depends less on how long you train than on how much the new task overlaps and conflicts with the old one.

Figure 11 · Drawn from the lesson's code

1 0 0 1 0 1 1 0 2 step of the second fine-tune (log scale) 0.0 0.2 0.4 0.6 0.8 1.0 held-out accuracy Fine-tuning on B after A: A collapses unless replayed accuracy on A (B only) accuracy on B (B only) accuracy on A (B + 10 replayed A) accuracy on B (B + 10 replayed A)

Fine-tuning on B after A: accuracy on A falls from 0.98 to near 0 within a few steps while B rises to 1; with 10 replayed A examples, A dips and then recovers to 0.945

Reading it: the x-axis is the step of the second fine-tune (log scale), the y-axis held-out accuracy. Solid lines are plain fine-tuning on B: B (green) climbs to 1 while A (blue) falls to about half in the same few steps, and A keeps sliding towards zero for as long as training continues. Dashed lines add 10 replayed A examples (section 3c): A still dips at first (to about 0.2 around step 40), then climbs back to 0.945 while B reaches 0.99.

Does a lower learning rate help here? Only in the sense of slowing the slide. At learning rate 0.02, or stopping after 10 steps, B reaches 0.975 and A still drops to 0.355.

Figure 12 · Drawn from the lesson's code

0.0 0.2 0.4 0.6 0.8 1.0 accuracy on the new task B 0.0 0.2 0.4 0.6 0.8 1.0 accuracy on the old task A Slower is not safer; replay escapes the trade-off A + B = 1 B only, lr 0.5 B only, lr 0.1 B only, lr 0.02 B + replay, lr 0.5

Accuracy on A against accuracy on B during the second fine-tune: every run on B alone traces the same curve whatever the learning rate and ends near zero on A, while the replay run climbs the right edge to the top-right corner

Reading it: each line traces one run: x is accuracy on B, y is accuracy on A, and each run starts at the top left (knows A, not B). The three learning rates (0.5, 0.1, 0.02) take very different numbers of steps but trace the same curve, close to the dotted line where A + B = 1: every point of B they gain costs about a point of A, and once B is learned, further steps slide A down the right edge towards zero. The replay run (dashed) follows the same curve at first, then climbs the right edge and ends in the top-right corner, knowing both. When the new data contradicts the old skill, slowing down can't help, because nothing in B's data says A still matters.

In code: sequential_run starts from the A fine-tune, trains on a new task (optionally with replay), and records both accuracies at every step; forgetting computes F.

Why it matters in practice. Real fine-tunes contradict the base model more often than you'd think: "always answer in JSON" contradicts "chat naturally"; "be terse" contradicts "explain in detail". Expect the old behaviour to vanish wherever your data overrides it, and test for it.

3c. Replay: keep practising the old skill

Everyday picture A pianist learning a new piece plays one old piece at the start of every practice session. It costs a few minutes and keeps the old repertoire alive.

Tiny worked example Add just 10 of task A's 200 training examples to B's 200: under 5% of the mix. The result: A 0.945, B 0.99 (up from A 0.025 without replay). Why can so few examples do so much? Look at the loss. Suppose that early in training the model already scores B well (loss 0.05 per example) but has started forgetting A (loss 3.0 on each replayed example):

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
examples of the new task 200
replayed examples of the old task 10
average loss on the new examples 0.05
average loss on the replayed examples 3.0
the loss training actually minimises: the average over every example in the mix 0.19

In words: "the loss on a mixed dataset is the average over all its examples, so each group counts in proportion to its size times its loss."

With the numbers: (200 × 0.05 + 10 × 3.0) / 210 = (10 + 30) / 210 = 0.19. The 10 replayed examples are under 5% of the data but contribute 0.143 of the 0.19: three quarters of the loss, and so most of the gradient. The old examples shout loudest exactly when they are being forgotten.

Level 3: in Python
n_new, L_new = 200, 0.05
n_old, L_old = 10, 3.0
# the replayed share of the data
round(n_old / (n_new + n_old), 3)  # → 0.048
# each group's contribution to the average loss
round(n_new * L_new / 210, 3), round(n_old * L_old / 210, 3)  # → (0.048, 0.143)
# L_mix
round((n_new * L_new + n_old * L_old) / (n_new + n_old), 2)  # → 0.19

Figure 16 · Diagram

Reading it: the old task's small sample joins the new data before training, so every gradient step sees both. This supplies exactly the counterweight that was missing in section 3a's diagram: when a change that helps B starts hurting A, the replayed examples' loss rises and pushes back.

In code: replay_mix appends the old examples, mixed_loss is the formula, and sequential_run with n_replay=10 runs the experiment.

Why it matters in practice. When fine-tuning a language model, mix some general instruction-following data into your task data, so the model keeps being a good assistant while it learns your task. When you can't replay (the old data is gone or private), methods such as elastic weight consolidation (Kirkpatrick et al.) instead penalise changes to the weights the old task relied on most.

The second fine-tune of 3b and the replay of 3c, in one lab.

Chapter 4

Overfitting a small dataset

Everyday picture A student with only 16 flashcards, three of which have the wrong answer on the back. For a while, studying teaches the pattern. Keep drilling and the student memorises every card word for word, including the three wrong answers, and gets worse on new questions. primer.ml.regularization builds this idea from scratch; here it is in a fine-tune.

Tiny worked example Fine-tune the base on only 16 examples of task A, 3 of them deliberately mislabelled, for 1,500 epochs. (With full-batch training, one step is one pass over the data, one epoch.)

At the best epoch (70) At the end (1,500 epochs)
training loss 0.373 0.005
validation loss 0.319 1.277

At the best epoch, training loss is higher than validation loss: the model is refusing to fit the three wrong labels, which is exactly right. By the end, training loss is nearly zero, so the wrong labels have been memorised, and validation loss has quadrupled.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the epoch 70, then 1,500
average loss on the 16 training examples 0.373, then 0.005
average loss on 200 held-out examples 0.319, then 1.277
the generalisation gap: how much worse the model does on data it hasn't seen −0.054, then 1.272

In words: "the gap is held-out loss minus training loss; a gap that keeps growing means the model is memorising rather than learning."

With the numbers: 0.319 − 0.373 = −0.054 at epoch 70; 1.277 − 0.005 = 1.272 at the end.

Level 3: in Python
L_train = {"best": 0.373, "end": 0.005}
L_val = {"best": 0.319, "end": 1.277}
# g = L_val − L_train at each point
{t: round(L_val[t] - L_train[t], 3) for t in L_val}  # → {'best': -0.054, 'end': 1.272}

Figure 17 · Drawn from the lesson's code

1 0 0 1 0 1 1 0 2 1 0 3 epoch (log scale) 0.0 0.5 1.0 1.5 2.0 2.5 loss A small dataset: learning, then memorising best epoch 70 training loss (16 examples, 3 mislabelled) validation loss (200 held-out examples)

Training loss falls steadily to near zero while validation loss bottoms out at epoch 70 and then climbs to four times its best

Reading it: the x-axis is the epoch (log scale), the y-axis the loss. Both curves fall at first while the model learns the real rule. At epoch 70 (the dashed line) validation loss bottoms out; after that the training curve keeps falling as the model memorises the three wrong labels, and the validation curve climbs. The dotted line is where early stopping with a patience of 20 epochs would end the run, keeping the weights from epoch 70.

Figure 18 · Diagram

Reading it: every epoch ends with a checkpoint and a held-out score. The loop keeps going while scores improve or patience remains, and when it stops, the checkpoint that ships is the marked best, not whatever the last epoch left behind.

In code: overfitting_run fine-tunes on the 16 examples and records both losses every epoch; primer.ml.regularization.early_stopping replays the validation curve and returns the best epoch and the stopping epoch.

Why it matters in practice. Fine-tuning datasets are small compared to pretraining, and big models memorise quickly, so fine-tunes typically run for only a few epochs. Save checkpoints, score each on the held-out set, and ship the best one.

The run above, with the epoch and the patience in your hands.

Chapter 5

Model merging

Move the three weights yourself first; section 5a then names what you did.

5a. Weight averaging and task arithmetic

Everyday picture Two editors each take a copy of the same draft and make tracked changes: one fixes the grammar, the other tightens the argument. You can apply both sets of changes to the original. If instead you "average" the two edited copies, each change is applied at half strength: half the grammar fixed, half the argument tightened.

Tiny worked example A three-weight base (1, 0, 2). One fine-tune moves it to (1.5, 0, 2); another to (1, −1, 2).

Weights Change from the base
base (1, 0, 2)
fine-tune on A (1.5, 0, 2) τ_A = (0.5, 0, 0)
fine-tune on C (1, −1, 2) τ_C = (0, −1, 0)
base + τ_A + τ_C (1.5, −1, 2) both changes in full
average of the two fine-tunes (1.25, −0.5, 2) both changes at half strength

The change a fine-tune made, θ_ft − θ_base, is its task vector. Adding task vectors to the base is task arithmetic.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
which task: a counter over the fine-tunes A, C
how many fine-tunes are merged 2
all the weights of the model fine-tuned on task (1.5, 0, 2)
the weights they all started from (1, 0, 2)
tau: task 's task vector, everything its fine-tune changed (0.5, 0, 0)
add up the task vectors τ_A + τ_C = (0.5, −1, 0)
lambda: how strongly to apply the combined changes 1
the merged model's weights (1.5, −1, 2)

In words: "each task vector is what its fine-tune changed; add the changes up, scale them by λ, and apply them to the base."

With the numbers: θ_merged = (1, 0, 2) + 1 × (0.5, −1, 0) = (1.5, −1, 2). With λ = 1/2: (1, 0, 2) + 0.5 × (0.5, −1, 0) = (1.25, −0.5, 2), which is exactly the average of the two fine-tunes. Averaging T fine-tunes is task arithmetic with λ = 1/T.

Level 3: in Python
theta_base = [1.0, 0.0, 2.0]
theta_A = [1.5, 0.0, 2.0]
theta_C = [1.0, -1.0, 2.0]
# τ_t = θ_t − θ_base
tau_A = [a - b for a, b in zip(theta_A, theta_base)]
tau_C = [c - b for c, b in zip(theta_C, theta_base)]
tau_A, tau_C  # → ([0.5, 0.0, 0.0], [0.0, -1.0, 0.0])
# θ_merged at λ = 1
[b + 1.0 * (a + c) for b, a, c in zip(theta_base, tau_A, tau_C)]  # → [1.5, -1.0, 2.0]
# λ = 1/2 ...
[b + 0.5 * (a + c) for b, a, c in zip(theta_base, tau_A, tau_C)]  # → [1.25, -0.5, 2.0]
# ... is the plain average of the two fine-tunes
[(a + c) / 2 for a, c in zip(theta_A, theta_C)]  # → [1.25, -0.5, 2.0]

Figure 21 · Diagram

Reading it: both fine-tunes start from the same base; that shared start is what makes their changes comparable. Subtracting the base turns each fine-tuned model into a task vector, the arrows are added and scaled, and the result is added back to the base. No data and no gradient steps are involved: merging is arithmetic on weights, and the merged model is the same size and speed as the base.

Figure 19 · Drawn from the lesson's code

0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 λ: strength of the summed task vectors 0.0 0.2 0.4 0.6 0.8 1.0 held-out accuracy of the merge base + λ(τ_A + τ_C): cosine(τ_A, τ_C) = -0.04 plain average (λ = 1/2) accuracy on A accuracy on C

Merging the fine-tunes on A and C: at lambda 1 the merged model scores 0.97 on both, while plain averaging (lambda one half) scores only 0.635 on A

Reading it: the x-axis is λ and the y-axis held-out accuracy of the merged model on A (blue) and on C (green); the dotted line is 0.5, a coin flip. At λ = 1/2, which is plain averaging, each skill is diluted: 0.635 on A. Around λ = 1 the merged model does both tasks at 0.97, as well as either specialist on its own task. Push λ further and the changes overshoot. Their task vectors are nearly perpendicular (cosine −0.04), so adding one barely disturbs the other.

In code: task_vector subtracts the base, merge adds scaled task vectors back, and merge_run merges two fine-tunes at several λ and scores the result.

Why it matters in practice. Merging combines skills trained separately, by different teams or on data that can't be pooled, without any further training and at no extra inference cost. Averaging several fine-tunes of the same task ("model soups", Wortsman et al.) often beats the best single one. Task vectors can also be subtracted: Ilharco et al. negated a task vector learned from toxic text to make a model less toxic.

5b. Interference: when task vectors collide

Everyday picture Two editors rewrote the same sentence in opposite directions, one making it warmer and one making it colder. Applying both sets of tracked changes gives a sentence neither intended.

Tiny worked example Our three-weight changes again, plus a new one: τ_A = (0.5, 0, 0), τ_C = (0, −1, 0) and τ_B = (−0.4, 0, 0.3). τ_A and τ_C touch different weights: no conflict. τ_B pulls the first weight the other way from τ_A. The cosine measures how aligned two changes are:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
dot product: multiply matching entries, then add 0.5 × (−0.4) = −0.2
a vector's length 0.5 and 0.5
the cosine of the angle between the two changes: 1 same direction, 0 unrelated, −1 opposite −0.8

In words: "multiply the changes weight by weight and add, then divide by both lengths, so only the direction counts."

With the numbers: τ_A · τ_B = 0.5 × (−0.4) + 0 + 0 = −0.2; ‖τ_A‖ = 0.5, ‖τ_B‖ = √(0.16 + 0.09) = 0.5; cos = −0.2 / 0.25 = −0.8: strongly opposed. cos(τ_A, τ_C) = 0: independent. See primer.ml.embeddings.similarity for the cosine from scratch.

Level 3: in Python
import math
tau_A = [0.5, 0.0, 0.0]
tau_B = [-0.4, 0.0, 0.3]
tau_C = [0.0, -1.0, 0.0]
def cos(u, v):
    dot = sum(a * b for a, b in zip(u, v))
    return dot / (math.sqrt(sum(a * a for a in u)) * math.sqrt(sum(b * b for b in v)))
# opposed changes
round(cos(tau_A, tau_B), 2)  # → -0.8
# changes to different weights
cos(tau_A, tau_C)  # → 0.0

In the toy, the fine-tunes on A and on B (which reverses A's rule) have task vectors with cosine −0.25, against −0.04 for A and C. Each fine-tune learned its rule everywhere, not just in its own region, so the two vectors rewrite the same weights in opposite directions.

Figure 20 · Drawn from the lesson's code

0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 λ: strength of the summed task vectors 0.0 0.2 0.4 0.6 0.8 1.0 held-out accuracy of the merge base + λ(τ_A + τ_B): cosine(τ_A, τ_B) = -0.25 plain average (λ = 1/2) accuracy on A accuracy on B

Merging the fine-tunes on A and B: at every lambda at least one task stays at or below a coin flip, and the best the merge manages on both at once is 0.525

Reading it: the same axes as the previous figure, now merging A with B. No value of λ lifts both lines: wherever one task improves, the other sits near or below a coin flip, and from λ = 1 on both stay there, because the two changes cancel. A model can do both tasks (replay in section 3c got 0.945 and 0.99), but this merge can't reach it: that model needs to tell the regions apart, and neither fine-tune learned to.

Figure 22 · Diagram

Reading it: a cheap check before merging is the cosine between task vectors. Near zero, the changes live in different weights and adding them usually works. Clearly negative, they fight: train on both datasets together if you can; if you can't, conflict-resolving merges such as TIES-merging (Yadav et al.) drop each task's smallest changes and, for each weight, keep only the changes that agree with the majority sign. Every path ends at the same place: score the merge on every task's held-out set.

In code: cosine measures the angle between two task vectors, and merge_run reports it alongside the merged accuracies.

Why it matters in practice. Merges are free to try, which makes them tempting to trust. A merged model can quietly lose a skill both parents had, so it is evaluated like any new model. The cosine tells you in advance which merges to be nervous about.

The two λ sweeps above, merged live and with λ in your hands.

Test yourself

9 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1When is fine-tuning the wrong tool, and what should you try first?Think it through, then reveal

When the model lacks facts, or the facts change: retrieval supplies them and is easy to update. When a clearer prompt with a few examples fixes the behaviour: that is cheaper and survives a base-model upgrade. Fine-tuning earns its cost when behaviour stays inconsistent under the best prompt, or when a long prompt sent millions of times costs more than the fine-tune.

Question 2Why build the held-out set before training, and why check it against the training set?Think it through, then reveal

So that no choice (prompt, learning rate, checkpoint) is made by looking at it, and it stays an honest measure. A training example that nearly copies a held-out one lets the model recite the answer, turning the eval into a memory test; deduplicating across the split prevents it.

Question 3A held-out set has 100 examples and the fine-tune scores 83% against the prompt's 80%. Has it won?Think it through, then reveal

Not yet. At 80% on 100 examples the 95% margin is about ±7.8 points, so a 3-point difference is well inside the noise. You need a larger held-out set (400 examples halve the margin) or a bigger difference.

Question 4If 10% of the eval's labels are wrong, what is the best score a perfect model can get?Think it through, then reveal

90%, because it is marked wrong on every mislabelled item. A 90%-accurate model would score 0.9 × 0.9 + 0.1 × 0.1 = 82%. Noisy labels shrink and blur the differences you are trying to measure.

Question 5What is catastrophic forgetting, and why does it happen?Think it through, then reveal

Training on a new task alone erodes, or wipes out, skills the model had. Every weight is shared between tasks, and only the new task's examples produce gradients, so nothing pushes back when a change that helps the new task hurts an old one. It grows with how far the weights move and with how much the new task conflicts with the old.

Question 6Lowering the learning rate did not stop task A being forgotten. Why not, and what works?Think it through, then reveal

Task B contradicts A on the same inputs, so any progress on B costs A; a lower rate only walks the same trade-off more slowly. Replay works: mixing even 5% of A's examples into B's data gives the model a reason to keep A, and those few examples carry most of the loss exactly when A is slipping.

Question 7Training loss keeps falling, but validation loss has risen since epoch 70. What is happening, and which checkpoint do you ship?Think it through, then reveal

The model has stopped learning the general rule and is memorising the training set, including its mislabelled examples. Ship the checkpoint from epoch 70, the best on the held-out set; early stopping automates exactly this.

Question 8What is a task vector, and why is averaging two fine-tunes the same as task arithmetic with λ = 1/2?Think it through, then reveal

A task vector is everything a fine-tune changed: θ_ft − θ_base. The average of two fine-tunes is (θ_base + τ_1 + θ_base + τ_2) / 2 = θ_base + ½(τ_1 + τ_2), which is task arithmetic with λ = 1/2, so each skill arrives at half strength.

Question 9When does merging fail, and how can you see it coming?Think it through, then reveal

When the task vectors change the same weights in opposite directions, so adding them cancels both skills. A clearly negative cosine between task vectors is the warning; the remedy is joint training or replay, or a conflict-resolving merge such as TIES, and a held-out check on every task either way.

Primary sources

The papers behind this lesson

Ilharco et al., Editing Models with Task Arithmetic (2022)

Defined task vectors as fine-tuned minus pretrained weights and showed that adding them combines skills, and negating them removes a behaviour.

Read the annotated companion →The paper ↗
Wortsman et al., Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time (2022)

Showed that averaging the weights of several fine-tunes of one base often beats the best single fine-tune.

The paper ↗
Yadav et al., TIES-Merging: Resolving Interference When Merging Models (2023)

Traced failed merges to small redundant changes and sign conflicts, and fixed both by trimming and electing a sign per weight.

The paper ↗
Kirkpatrick et al., Overcoming catastrophic forgetting in neural networks (2017)

Introduced elastic weight consolidation, which slows learning on the weights most important to earlier tasks.

The paper ↗
Biderman et al., LoRA Learns Less and Forgets Less (2024)

Measured on real language models that LoRA keeps more of the base model's abilities than full fine-tuning, at the price of learning the new task less completely.

The paper ↗
Zhou et al., LIMA: Less Is More for Alignment (2023)

Fine-tuned a large base model on 1,000 carefully curated examples and got a strong assistant, evidence that example quality matters more than quantity.

Read the annotated companion →The paper ↗
Lee et al., Deduplicating Training Data Makes Language Models Better (2021)

Found widespread near-duplicates in standard datasets, including between training and test sets, and showed that removing them reduces memorisation.

The paper ↗

Researcher's shelf

Further reading

  • Goodfellow et al., An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks (2013): https://arxiv.org/abs/1312.6211
  • Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (2021): https://arxiv.org/abs/2106.09685
  • Ilharco et al., Editing Models with Task Arithmetic (2022): https://arxiv.org/abs/2212.04089
  • Yadav et al., TIES-Merging (2023): https://arxiv.org/abs/2306.01708
  • Hugging Face TRL, supervised fine-tuning trainer: https://huggingface.co/docs/trl/sft_trainer
  • Hugging Face PEFT, parameter-efficient fine-tuning (LoRA and friends): https://huggingface.co/docs/peft/index
  • mergekit, an open-source toolkit for merging models: https://github.com/arcee-ai/mergekit

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.