Prompt chaining: break one big prompt into steps
Table of contents
Here is a prompt a lot of people write on their first hard task: "Read these twelve pages of meeting notes, pull out the decisions, figure out which ones affect the roadmap, rank them by risk, and write me a one-paragraph update for my manager." One prompt, five jobs. The model does all five at once, badly. The decisions are half-right, the ranking is a guess, and the paragraph blends everything into a smooth summary that is confidently wrong in two places.
The fix is not a better mega-prompt. It is to stop asking for five things in one breath. Ask for the decisions. Take that output, and ask which ones touch the roadmap. Take that, and ask for the ranking. Then write the paragraph from the ranked list. Each step is small enough that the model can actually do it, and you can see the output of each one before it feeds the next.
That is prompt chaining, and it's the difference between a model that feels unreliable and one that feels like a tool. This piece covers what prompt chaining is, a chain you can copy, when to reach for it instead of a single prompt or few-shot examples, where chains break, and why a chain you save and reuse has an expiry date you will not see.
What is prompt chaining?
Prompt chaining is a prompt engineering technique that breaks a complex task into a sequence of smaller prompts, where the output of each prompt becomes the input to the next. Instead of one instruction that tries to do everything, you get a pipeline: focused step, checked output, focused step, checked output, final result.
The reason it works is not magic, it is attention. A model asked to do one thing spends all of its capacity on that thing. A model asked to do four things at once splits its focus and cuts corners on all four. The Prompt Engineering Guide describes chaining as breaking a task into subtasks and prompting the model with each in turn, using the response to the first as part of the input to the next. IBM's write-up on prompt chaining frames it the same way, as generating a final output by following a series of prompts rather than a single one. Anthropic's guide to building effective agents lists prompt chaining as the first and simplest workflow pattern: decompose a task into a fixed sequence of steps, and optionally check the output between them before continuing.
There is a second, quieter benefit. When a single prompt fails, you have no idea which part went wrong, so you rewrite the whole thing and hope. When a chain fails, you can see exactly which step produced the bad output and fix that one link. The chain is debuggable in a way the mega-prompt never is. If you already treat your notes and instructions as things you maintain rather than write once and forget, a chain fits the same habit: it's a workflow you can inspect a step at a time.
A chain you can copy
The whole technique is easier to show than to explain. Here is a three-step chain for the meeting-notes task from the opening. Run each prompt, paste its output into the next, and read what comes back at each stage.
Step 1: extract, do not summarize.
Below are my raw meeting notes. Extract every decision that was
actually made, one per line, in the format:
DECISION: <what was decided> | OWNER: <who> | DATE: <if stated>
Do not include discussion, opinions, or things that were only
proposed. Only decisions that were settled.
NOTES:
[paste raw notes]Step 2: filter against a goal. Feed Step 1's list in.
Here is a list of decisions from a meeting. Keep only the ones that
change our Q3 roadmap. For each, add one line: WHY IT MATTERS.
Drop the rest silently.
DECISIONS:
[paste Step 1 output]Step 3: write from the filtered list. Feed Step 2's output in.
Write a 4-sentence update for my engineering manager based only on
the decisions below. Lead with the highest-risk item. Plain language,
no hype, no filler. Do not invent anything not in the list.
DECISIONS:
[paste Step 2 output]Notice what each step buys you. Step 1 forces the model to separate decisions from noise, which is the part it gets wrong when it is also trying to write prose. Step 2 does the judgment call in isolation, so you can eyeball the filter before it colors the writing. Step 3 writes from a clean, short input, so it has nothing to blur. If the final paragraph is wrong, you know within seconds whether the fault is extraction, filtering, or writing, because you kept all three outputs.
You do not need a framework for this. A chain is just running prompts in order and pasting outputs forward. Tools like LangChain automate the plumbing when a chain runs often enough to be worth wiring up, but the technique itself is copy, paste, read, repeat. Start there, and reach for automation only when the manual version has already earned its keep.
When to chain, and when not to
Chaining is not always the right move. It costs more calls, more time, and more places to go wrong, so a task that a single prompt handles well should stay a single prompt. The question is which shape fits the job. This is also the question the search results for prompt chaining keep dancing around, because the honest answer is a comparison, not a definition.
| Situation | Reach for | Why |
|---|---|---|
| One clear task, one output | A single prompt | Chaining adds cost and failure points for no gain. |
| A task with distinct stages that build on each other | Prompt chaining | Each stage gets full attention, and you can inspect the handoffs. |
| The model keeps getting the format or category wrong | Few-shot prompting | A few worked examples fix format and labeling faster than more steps. |
| One hard reasoning problem, solved in place | Chain-of-thought | Ask the model to reason step by step inside one prompt; the steps are its thinking, not separate calls. |
| A standing rule that should apply to everything | The system prompt | Durable instructions belong above the task, not repeated in every chain. |
The two that get confused most are prompt chaining and chain-of-thought, because both involve "steps." The difference is who does the stepping. Chain-of-thought is one prompt where the model reasons through steps out loud before answering; the steps are internal to a single call. Prompt chaining is several prompts where you own the boundaries, see each intermediate output, and decide what feeds the next. Use chain-of-thought for a single hard problem. Use chaining when the stages are genuinely different jobs, or when you need to check the work partway through.
Few-shot is the other near neighbor. If the model's problem is that it keeps producing the wrong shape, adding examples usually fixes it more cheaply than adding a step. Chaining and few-shot are not rivals, though: the best version of Step 1 above might carry two examples of a well-extracted decision. Reach for examples when the failure is format or category, and reach for a new link in the chain when the failure is that one prompt is doing too much.
Where chains break
A chain trades one big failure for several small ones, and most of the small ones are quiet. Knowing the shapes in advance is the difference between a chain you trust and one that fails without telling you.
| Failure mode | What it looks like | What to do |
|---|---|---|
| Error compounding | A small mistake in Step 1 is treated as fact by Steps 2 and 3 and amplifies. | Check the output of early steps, not just the final one. The first link matters most. |
| Context drift | Details get dropped or subtly reworded at each handoff, so the final output has drifted from the source. | Carry the original input forward where accuracy matters, not just the previous step's output. |
| First-prompt dependency | The whole chain's quality is capped by Step 1; a weak extraction cannot be rescued downstream. | Spend most of your effort on the first prompt. It is load-bearing. |
| Cost and latency multiplication | Three prompts is roughly three times the calls, tokens, and wait. | Only chain when the accuracy gain is worth the multiplier. A good single prompt beats a needless chain. |
| Silent staleness | A saved chain keeps running long after the task, format, or facts it assumed have changed. | Re-read the chain on a schedule. This one gets its own section. |
The first three are the ones people underestimate. Because each step looks reasonable in isolation, a chain can produce a polished final answer that is wrong for a reason buried two steps back. The habit that catches it is boring and effective: keep the intermediate outputs, and read the early ones first when something looks off. The same pressure shows up inside a single long context as context rot, where the model's accuracy slides as the window fills, so a shorter chain that keeps each step's input lean is easier to keep honest than one giant prompt. The same discipline that keeps a note honest, checking the source rather than trusting the summary, is what keeps a chain honest.
A saved chain is an artifact that rots
The first time a chain works, you save it, because rebuilding it every week is tedious. That is the right instinct and the start of the problem. A chain is a set of instructions, and instructions bake in the world as it was the day you wrote them.
The meeting-notes chain assumed a manager who wanted four sentences, a roadmap called Q3, and decisions that look a certain way. When the manager changes, or the quarter rolls, or the team starts recording decisions in a different format, the chain does not know. It keeps extracting, filtering, and writing exactly as instructed, producing output that is well-formatted and quietly wrong. Nothing errors. This is knowledge decay, the same rot that turns a good system prompt into scar tissue and a good runbook into a liability, sitting one level up in a workflow you stopped looking at.
The parts that go stale fastest are the ones tied to something that changes on its own schedule:
- A baked fact. A step that says "our pricing is X" or "the API returns Y" carries an expiry date in a place with no expiry field.
- A named target. "Write for my manager Dana" is wrong the week the org chart moves, and the chain cannot tell.
- A model assumption. A chain tuned around one model's quirks can misfire on the next one, and switching models freely is the whole point of being able to bring your own provider.
When the context a chain leans on comes from your own notes, the fix is upstream: the notes have to stay true, or every workflow built on them inherits the error. That is the premise of an assistant that treats your knowledge as something to maintain rather than only store. Scribelet's AI memory exists so the facts your AI works from can be listed, checked, and corrected in one place, instead of frozen into a chain where they silently age. A chain, like a generated prompt, is only trustworthy on the day you last read it.
Chain when the task earns it
Prompt chaining is the most useful prompt technique you can learn in an afternoon, and the case for it is plain: a model does one thing well and four things poorly, so give it one thing at a time. Break the task where the jobs are genuinely different. Keep the intermediate outputs so you can see which link failed. Spend your effort on the first prompt, because everything downstream inherits it.
And treat a chain you save the way you would treat any note you rely on: as something that was true when you wrote it and needs a look before you trust it again. The point of chaining is not to build a machine you never touch. It's to break a hard task into pieces small enough to check, and then to keep checking them. Set up your first desk and see what it looks like when an AI has to show its work at every step, not just hand you a confident final answer.
Share this article