Skip to content

Few-shot prompting: when examples help and when they hurt

Scribelet Team
10 min read

A team ships a few-shot prompt to sort support tickets: three examples labeled urgent, normal, and low. It works. Six weeks later the same prompt is calling half of everything urgent, and nobody changed a line.

What changed was the tickets. The three examples were written back when "urgent" meant a payment outage. The new tickets are feature requests and billing questions, and the model, still matching against three months-old examples, keeps forcing them into a shape that no longer fits.

That is few-shot prompting working exactly as designed and failing anyway. The examples did their job. They just encoded a world that moved.

This piece covers what few-shot prompting is, when to use no examples versus one versus a handful, a prompt you can copy, the failure modes the tutorials skip, and why the examples you paste today have an expiry date you will not see.

What is few-shot prompting?

Few-shot prompting is a technique where you put a small number of examples, usually two to five, of the input you will give and the output you want directly inside the prompt, before the real request. The model reads the pattern in those examples and applies it to the new input, with no retraining.

The name comes from the paper that made the technique famous. When OpenAI introduced GPT-3 in 2020, the title was Language Models are Few-Shot Learners: the finding was that a large enough model could learn a task from a few examples in the prompt, with no gradient updates and no fine-tuning. The mechanism has a name, in-context learning. The model's weights never change; it conditions its next answer on the examples sitting in front of it. IBM's write-up on zero-shot versus few-shot prompting frames it the same way: the examples are guidance, not training.

That is the whole idea, and it is genuinely useful. An example is worth a paragraph of instructions. "Return the answer as JSON with keys name and role" is a rule the model might follow. One filled-in example of that JSON is a rule it can see.

The catch is that an example carries more than you meant to put in it. Its format, its wording, its order, and the mix of cases you happened to include are all instructions, whether you intended them or not.

Zero-shot, one-shot, few-shot: which to reach for

The "shot" is just the count of examples. Zero-shot is no examples, only an instruction. One-shot is a single example. Few-shot is two or more. Each is the right tool in a different place, and reaching for the wrong one is the first mistake.

ApproachExamplesReach for it whenThe risk
Zero-shotNoneThe task is common and the format is simple. Modern models already know "summarize this" or "translate to French".The model guesses the format, and the guess drifts between calls.
One-shotOneYou need a specific output shape and one example pins it: a layout, a tone, a particular JSON.A single example reads as the answer, not a pattern. The model over-copies it.
Few-shotTwo to fiveThe task has edge cases, categories, or a judgment call one example cannot show. Classification, extraction with exceptions.More examples, more tokens on every call, plus the ordering bias below.
Fine-tuningDozens to thousandsYou run the same task at volume and the examples no longer fit in the prompt.A trained model is a frozen artifact too, and a costlier one to change.

Diagram of the zero-shot, one-shot, few-shot spectrum and what each one trades off

The honest default for most one-off work is zero-shot first. Try the instruction alone. Add examples only when the output is wrong in a way an example would fix, which is almost always a format or edge-case problem, not a knowledge problem. If the model does not know something, examples will not teach it. They will only make it guess more confidently, which is the opposite of what you want when the thing you are prompting has to be checked against a source afterward. And when the failure is not a missing example but a single prompt trying to do too much at once, the fix is not more examples; it is breaking the task into a chain of smaller prompts, each doing one job.

A few-shot prompt that works

Here is a few-shot prompt for a real task: turning a messy meeting note into structured action items. The shape matters more than the words.

Extract action items from meeting notes. For each, return the owner,
the task, and a due date if one is stated. If no owner is named, use "unassigned".
 
Input: Sarah will send the vendor contract by Friday. We still need
someone to review the Q3 numbers.
Output:
- owner: Sarah | task: send the vendor contract | due: Friday
- owner: unassigned | task: review the Q3 numbers | due: none
 
Input: Mike is going to update the runbook after the migration.
Let's revisit pricing next week.
Output:
- owner: Mike | task: update the runbook | due: after the migration
- owner: unassigned | task: revisit pricing | due: next week
 
Input: {your meeting note here}
Output:

Three things make this work, and each is a lever:

  • Consistent structure. Every example uses the same labels (owner, task, due) in the same order. Models copy structure faster than content, so a wobble in your examples becomes a wobble in the output.
  • An edge case on purpose. The "unassigned" and "none" cases appear in the examples, so the model has seen what to do when a field is missing. An example set that only shows the happy path teaches the model to invent a value rather than leave it blank.
  • The pattern is the point. Two examples, not eight. Once the shape is unambiguous, more examples add token cost and the ordering bias below, not accuracy.

If you use an assistant that keeps a system prompt or custom instructions, this is exactly the kind of thing that should not live there permanently. A worked example is paid for on every request and relevant to few of them, which is why long examples belong in the request that needs them, not in the standing brief.

How few-shot prompting quietly breaks

The tutorials stop at "use consistent formatting". The failures that actually cost you are more specific, and most stay hidden because the output reads fluently while it goes wrong.

Failure modeWhat you seeWhat caused it
Recency biasThe model favors the label or format of your last exampleModels weight the most recent example most heavily. Put the most typical case last, not an edge case.
Majority-label biasWith three "urgent" examples and one "normal", everything trends urgentThe label mix in your examples leaks into the output. Balance the classes, or the model learns your ratio instead of the task.
Format lock-inA new input that does not fit the example shape gets forced into it anywayThe examples taught a rigid template, and the model bends the input rather than break the pattern.
Example leakageNames or values from your examples show up in answers about unrelated inputsThe model treats example content as context, not just as a pattern. Use obviously generic, fake values.
Domain overfittingGreat on inputs like your examples, poor on everything elseYour examples all came from one domain, so the model narrowed to it. Cover the range you actually expect.
Stale examplesOutput was right for months, then slowly went wrong with no code changeThe examples encode a format, taxonomy, or tone that has since moved. See below.

The first two have a name in the research. The paper Calibrate Before Use showed that few-shot accuracy swings widely based on the order of the examples, the specific examples chosen, and the label distribution, and that the model leans toward answers it saw recently and often. The practical version is short: your examples are not neutral. Their order and their balance are instructions you did not know you were writing.

Your examples have a shelf life

The last row of that table is the one nobody plans for, and it is the one worth dwelling on.

An example is a snapshot. When you paste two support tickets labeled by your 2025 taxonomy, or a meeting note formatted the way your team wrote them last quarter, you are freezing a moment. The prompt keeps steering toward that moment long after the world it captured has changed.

The classifier from the opening did not break because the model got worse. It broke because "urgent" drifted while the three examples still defined it the old way. The output stayed confident the whole time, which is exactly why it took six weeks to notice.

Loop diagram: an example pins a format, the real format drifts, the model keeps matching the stale example, and the output stays confidently wrong

This is the same failure that turns a good system prompt into scar tissue and a good runbook into a liability. It is knowledge decay, one layer down: the examples are notes, the model reads them on every call, and nothing ever renders them visibly wrong.

The examples that rot fastest are the ones tied to something that changes on its own schedule:

  • A taxonomy. Category labels, priority levels, tag names. When the set changes, every example using the old set is teaching the wrong map.
  • A format. A JSON schema, a date convention, a field that got renamed. The examples keep producing the old shape.
  • A fact. An example modeling "our current plan is X" bakes a fact with an expiry into a place that has no expiry field.

If the examples you feed a model come from your own notes, the fix is upstream: the notes have to stay true. That is the premise of an assistant that treats your knowledge as something to maintain rather than just store. Scribelet's AI memory and background verification exist so the material your AI works from can be listed, checked, and corrected in one place, instead of copied into a prompt where it silently ages.

When to stop adding examples

There is a point where the answer is not another example. Few-shot has a ceiling, and pushing past it spends tokens for worse results.

Stop adding examples when:

  • You are past five or six and still not happy. The gain from each example flattens fast. If six have not fixed it, the problem is the instruction or the task, not the count. When it is the instruction, having the model draft the prompt for you beats reaching for a seventh example.
  • The examples no longer fit the context window. Long few-shot prompts get expensive on every call, and the earliest examples start getting ignored anyway. If you need dozens, you have outgrown the technique.
  • The task needs knowledge, not a pattern. Examples show the model how to answer, not what is true. For facts, reach for retrieval, not more examples.

At volume, the honest next step is fine-tuning: train the pattern into the model so it does not cost prompt space on every call. But note the trade. A fine-tuned model is a frozen artifact, and a costlier one to change than a prompt. The examples that rot in a few-shot prompt at least sit in a text file you can read; the same drift baked into fine-tuning weights is invisible and costs a training run to fix. When you choose your own provider and model, that trade is at least yours to make on purpose.

Examples are an artifact, not a setting

Few-shot prompting gets treated as a trick: paste a couple of examples, get better output, move on. It works well enough often enough that the examples turn permanent the day they are added, and then they sit there, defining a task by a handful of cases from one particular week.

They deserve the same care as anything else the machine reads on every request. Keep them few. Keep them balanced. Keep them in version control, so a change is a diff and not a mystery. And re-read them on a schedule, because the examples that were perfect in March are quietly teaching your model the wrong lesson by September, and the output will never tell you.

If the examples your AI learns from come from your own notes, they should be as inspectable and as maintainable as anything else you write down. Set up your first desk and see what it looks like when an AI has to show what it knows and where it learned it.

Share this article

We use cookies for analytics to improve your experience. Learn more