Spec-driven development: write the spec your agent follows
Table of contents
Here is a way to build a feature with an AI coding agent that almost works. You open the chat and type: "Add a way for users to export their data as CSV, with a button in settings, and make sure it handles large accounts." The agent thinks for a while and produces four hundred lines across six files. Some of it is right. The button is in the wrong place, "large accounts" turned into an arbitrary 10,000-row cap you never asked for, and the CSV escaping breaks on the first name with a comma in it. You did not ask for any of that. You also did not say not to.
The problem is not the model. The problem is that you handed a paragraph to something that will happily fill every gap you left with a guess. A paragraph has a lot of gaps. The fix is the oldest idea in software, arriving in a new place: write down what you actually want before anything gets built, in enough detail that the gaps are closed.
That's spec-driven development, one discipline inside the broader practice of agentic coding, and it's the fastest-growing answer to a problem AI coding created. This piece covers what spec-driven development is, a spec you can copy and adapt today, when to reach for it instead of just prompting or writing tests first, where the approach breaks, and why the spec you write this week is quietly wrong by next month if you never look at it again.
What is spec-driven development?
Spec-driven development (SDD) is a workflow where you write a structured, human-readable specification of what to build, review and agree on it, and then have an AI coding agent implement against it, rather than prompting the agent directly and correcting it after the fact. The spec is the source of truth. The code is derived from it.
The shape is consistent across the people defining it. Martin Fowler's team describes SDD as writing a well-considered spec first, then using it in the AI-assisted workflow as the durable artifact the code answers to. GitHub, which shipped an open-source toolkit called Spec Kit for exactly this, frames the loop as capturing intent in a spec, planning from it, then generating code inside those constraints. IBM's write-up on spec-driven development puts it plainly: a detailed specification is authored and agreed upon before implementation begins.
What all of them are reacting to is "vibe coding," the prompt-and-pray style where you describe a feature in a sentence and accept whatever comes back. Vibe coding is fine for a throwaway script. It falls apart the moment the thing has to be correct, because the model is filling your gaps with statistically plausible choices, not the choices you would have made. A spec is how you make those choices yourself, up front, once, instead of catching them in review one wrong guess at a time.
If you have ever written a software design document or a technical spec, you already know most of this. The document is the same idea. What changed is the reader. It used to be a teammate who would ask you questions when something was unclear. Now it is an agent that will not ask. It will guess, and it will guess confidently, so the spec has to close the gaps a human colleague would have flagged.
A spec you can copy
The technique is easier to show than to argue about. Here is a spec skeleton for the CSV export from the opening. It's deliberately short. A spec is not a novel; it's the set of decisions you don't want the agent making for you.
# Feature: CSV export of user data
## Goal
Let a signed-in user download all their own records as a CSV file.
## User-facing behavior
- A "Export CSV" button on the Settings > Data page, under "Your data".
- Clicking it downloads a file named `export-<YYYY-MM-DD>.csv` immediately.
- No email, no background job, no "we'll send you a link".
## Scope
- Export ONLY the current user's own rows. Never another user's data.
- Include these columns, in this order: id, title, created_at, status.
- Exclude soft-deleted rows (deleted_at is not null).
## Rules and edge cases
- Escape any field containing a comma, quote, or newline per RFC 4180.
- created_at in ISO 8601 UTC.
- An account with zero rows downloads a header-only file, not an error.
- Do NOT add a row limit. Stream the response if the result is large.
## Out of scope
- No XLSX, no PDF, no column picker. CSV only, all columns, this release.
## Done when
- A user with 0, 1, and 50,000 rows all get a correct file.
- A title containing `a,"b"` round-trips correctly through a spreadsheet.Read what that buys you. Every wrong guess from the opening is now a line the agent cannot get wrong: the button's location, the "large accounts" behavior, the escaping rule, the filename. The "Out of scope" section is doing as much work as the rest, because it stops the agent from helpfully building three formats you did not ask for. The "Done when" section gives it something to check its own work against, which is the part vibe coding never has.
You hand that spec to the agent as the task. The agent implements it. Then you check the result against the spec, not against a vague memory of what you wanted, because the spec wrote the memory down. When something is wrong, you fix the spec first and regenerate, so the spec stays the source of truth instead of drifting behind the code on day one.
You do not need a framework to start. A spec is a markdown file and some discipline. Tools like Spec Kit or Kiro add structure, templates, and a repeatable command loop when you are doing this every day, but the technique underneath is exactly what you just read: decide, write it down, then build.
When to reach for a spec, and when not to
Spec-driven development is not free. Writing the spec is real work, and for a small enough task it is slower than just asking. The skill is knowing which shape fits the job. This is also the question the search results dance around, because the useful answer is a comparison, not a definition.
| Situation | Reach for | Why |
|---|---|---|
| A throwaway script or a one-line change | A single prompt | Writing a spec costs more than the fix. Just ask. |
| A feature with real behavior, edge cases, or "don't do X" rules | Spec-driven development | The spec closes the gaps the agent would otherwise guess. |
| You care most about verified behavior, not structure | Test-driven development | Write the failing tests first; the tests are the spec. Pairs well with SDD. |
| A cross-team product decision about what to build | A PRD or design doc | The audience is people, and the questions are product, not implementation. |
| A recurring instruction for every task in a repo | The system prompt or an agent rules file | Durable, cross-cutting rules belong above the task, not in each spec. |
The pairing worth noticing is spec-driven and test-driven development, which people treat as rivals and are not. A spec says what to build in prose; tests say what "correct" means in code. The "Done when" section of the spec above is a test suite waiting to be written. The strongest version of this workflow writes the spec, turns the "Done when" bullets into failing tests, and lets the agent code until the tests pass against the spec. SDD sets the intent; TDD verifies it.
The other honest boundary is the design doc. A software design document answers "should we build this, and how does it fit the system," for an audience of humans who will push back. A spec answers "build exactly this," for an agent that will not. The same project often has both: the design doc decides the shape, the spec hands the agent the details. Do not collapse them into one document that serves neither reader well.
Where spec-driven development breaks
A spec trades the risk of a wrong guess for a set of quieter risks, and the quiet ones are the ones that cost you later. Knowing them in advance is the difference between a spec that helps and a spec that lies to you.
| Failure mode | What it looks like | What to do |
|---|---|---|
| Over-speccing | You spend an hour specifying a ten-minute change, or dictate implementation details the agent should own. | Spec the decisions and constraints, not the code. If a line does not close a gap, cut it. |
| Under-speccing | The spec reads clean but leaves the same gaps as a prompt, so the agent still guesses. | Write the edge cases and the out-of-scope list. Those are where guesses hide. |
| Spec drift | The code changes in review or later PRs, the spec does not, and the two quietly diverge. | Fix the spec first and regenerate, or delete it when it stops being true. A stale spec is worse than none. |
| Trusting a passing agent | The agent says it matched the spec; it matched its reading of the spec. | Verify against the spec yourself, especially the "Done when" cases. The spec is the check, not the agent. |
| Baked-in facts | The spec hardcodes an API shape, a limit, or a dependency that changes on its own schedule. | Reference the source of that fact instead of copying it, and re-read specs that name external systems. |
The first two are about writing the spec well, and they trade off against each other: over-spec and it is slower than prompting, under-spec and it is prompting with extra steps. The line is the gap. A good spec line closes a decision the agent would otherwise make wrong. A wasted line describes something obvious or something the agent should decide on its own.
The last three are the ones that outlive the pull request, and they share a cause. A spec is a document, and documents freeze the world as it was the day you wrote them. The same discipline that keeps a note honest, checking the source instead of trusting the copy, is what keeps a spec honest past its first use.
A spec is an artifact that rots
Here is the part almost nobody writing about spec-driven development mentions. The first time a spec works, you keep it. You put it in the repo next to the code, because it explains the feature and you might need it again. That's the right instinct and the start of a slow problem.
The moment the code changes and the spec does not, the spec starts lying. Someone tweaks the export to include a fifth column in a later PR, the "columns, in this order" line still says four, and now the most authoritative-looking document about that feature is wrong. The next person to read it, or the next agent asked to modify the feature "according to the spec," inherits the error. Nothing broke. Nothing warned anyone. The spec just aged out of agreement with reality while sitting still, which is exactly how outdated documentation happens one file at a time.
This is knowledge decay, and a spec is unusually exposed to it because it is precise. A vague doc is hard to falsify. A spec that says "these four columns, in this order" is trivially falsifiable and therefore trivially wrong the day a fifth column ships. The precision that makes a spec useful to an agent is the same precision that makes it stale fast. The parts that rot first are the ones tied to something with its own clock:
- A hardcoded external fact. "The payments API returns
status: paid" carries an expiry date the day that API changes, and the spec has no expiry field. - A reference to code that moved. A spec that names files, functions, or limits is wrong the refactor after it was written.
- An out-of-scope line that quietly became in scope. "No XLSX this release" is true for exactly one release.
A spec you regenerate from and then update is a living artifact. A spec you write once and leave in the repo is a trap with good formatting. The workflow only pays off if the spec stays true, which means someone, or something, has to notice when it stops being true. That is the premise Scribelet is built on: your notes, docs, and the facts your AI works from are things to maintain, not just store. When the reference material an agent leans on lives somewhere that can flag its own staleness, a drifting spec gets caught instead of quietly shipping the wrong column. A spec, like the memory an AI keeps, is only trustworthy on the day you last checked it.
Write the spec, then keep it honest
Spec-driven development is the correction AI coding needed. A model will fill every gap you leave with a confident guess, so leave fewer gaps: decide the behavior, the edge cases, and the out-of-scope list before anything gets built, and hand the agent a spec instead of a paragraph. Keep it short enough to be worth writing and specific enough to close the guesses.
Then treat the spec the way you would treat any note you rely on: as something that was true when you wrote it and needs a look before you trust it again. The point of a spec is not a document you file and forget. It is a decision you made on purpose, written down so an agent cannot undo it, and worth keeping true for as long as the feature lives. Set up your first desk and see what it looks like when the source of truth is something that tells you when it has gone stale, instead of aging in silence.
Share this article