Agentic coding: how to work with an AI coding agent
Table of contents
The job changed and nobody sent a memo. A year ago, writing software meant typing it. Today a growing number of developers spend their day describing a task, watching an agent read files, run commands, and edit code for ten minutes on its own, and then reviewing what came back. The typing is increasingly the machine's job. The deciding, framing, and checking are yours.
That shift has a name. Agentic coding is the practice of handing a unit of work to an AI agent that can plan the steps, edit files, run tests, and iterate against the result, rather than autocompleting a line at a time or pasting snippets out of a chat window. It is the difference between an assistant that finishes your sentence and a colleague you can give a ticket to. The productivity stories are real, and so are the ways it goes quietly wrong.
This piece is the practical version: what agentic coding actually is, a workflow you can copy today, when to reach for an agent and when to keep your hands on the keyboard, where the approach breaks, and the failure almost nobody names, which is that the context you feed the agent decays the same way every other document does.
What is agentic coding?
Agentic coding is a development approach where an autonomous AI agent takes an active role in writing, testing, and modifying software, using tools like reading and writing files, running shell commands, and searching the web, instead of only suggesting text for you to accept. IBM's write-up on agentic coding frames it the same way: coding agents combine the reasoning of a language model with access to real developer tools and the ability to execute.
The word doing the work is "agent." A plain code assistant predicts the next token. An agent runs a loop: it forms a plan, takes an action, observes the result, and decides the next action, repeating until the task is done or it gets stuck. That loop is why an agent can open a failing test, read the traceback, edit three files, rerun the test, and try again without you in the seat for each step.
It helps to place agentic coding against its neighbors, because the terms blur together in practice:
- Autocomplete finishes the line you are typing. You drive; the tool guesses ahead.
- Chat-assisted coding answers a question or writes a snippet you paste in. You are still the one editing files.
- Agentic coding takes a whole task and does the editing, running, and checking itself. You frame and review.
- Vibe coding is agentic coding with the review turned off: you accept whatever comes back without reading it, which is fine for a throwaway and a liability for anything that has to be correct.
The distinction that matters is where the human sits. In agentic coding you move up the stack, from writing the code to specifying and verifying it. Do that well and you get leverage. Do it carelessly and you get a large volume of plausible code nobody actually understands.
A workflow you can copy
The single biggest predictor of whether agentic coding helps or hurts is how you hand off the work. A one-line prompt gets you a confident guess. A framed task gets you something you can check. Here is the loop that practitioners keep converging on, distilled from write-ups like Drew Breunig's 10 lessons for agentic coding and Armin Ronacher's agentic coding recommendations.
The frame is the part you write. Everything else is the agent's job, and your review is the gate at the end. A frame is not a paragraph; it is the set of decisions you do not want the agent making for you. Give it to the agent as the task:
## Task
Add rate limiting to the public API.
## Context the agent should read first
- The middleware lives in src/server/middleware/.
- We already use Redis (see src/lib/redis.ts). Reuse that client.
- Follow the error format in src/server/errors.ts.
## What "done" means
- 100 requests per minute per API key, sliding window.
- Over the limit returns HTTP 429 with a Retry-After header.
- Unauthenticated requests are limited by IP, same numbers.
## Constraints
- Do NOT add a new dependency; use the existing Redis client.
- Do NOT change the public response shape of any existing endpoint.
## When you are done
- Add tests for: under limit, at limit, over limit, and IP fallback.
- Run the full test suite and show me the output before finishing.Read what that frame prevents. The agent will not invent a new rate-limiting library, will not pick 60 requests instead of 100, and will not skip the IP fallback you care about, because each of those is now a line instead of a guess. The "context the agent should read first" section is doing quiet, heavy work: it points the agent at the code that already exists so it extends your system instead of building a parallel one. The "when you are done" section gives the agent something to check itself against, which is the part a bare prompt never has.
Two habits make the loop pay off. First, keep tasks small enough to review. An agent can write a thousand lines in a minute; you cannot honestly review a thousand lines in a minute, and the review is the whole point. Second, keep the agent's standing context in the repo, not in your head. Most coding agents read a project file (often named AGENTS.md or similar) on every run for durable rules like "run the linter before finishing" or "we use pnpm, not npm." That file is the agent's system prompt for your codebase, and it is where cross-cutting instructions belong so you do not repeat them in every task.
When to reach for an agent, and when not to
Agentic coding is a tool, not a religion. Some work suits an agent perfectly; some is faster and safer by hand. The skill is telling them apart before you have burned twenty minutes finding out. This is the question the definitions skip, because the useful answer is a comparison.
| Situation | Reach for | Why |
|---|---|---|
| A well-scoped task with clear "done" criteria and existing patterns to follow | An agent | This is the sweet spot. Frame it, delegate, review the diff and the tests. |
| A one-character fix or something you can type faster than you can explain | Your keyboard | Framing the task costs more than the fix. Just do it. |
| Unfamiliar code you need to understand, not just change | Chat, then decide | Use the agent to explain and explore first. Delegating a change you cannot review is how bugs ship. |
| A large, cross-cutting change with real design questions | A spec first, then an agent | Decide the shape yourself in a spec, then let the agent implement against it. |
| Anything touching security, money, or data you cannot easily undo | Your keyboard, or an agent on a very short leash | The cost of a confident wrong guess is too high. Review every line, or write it yourself. |
The pattern under the table: delegate when the work is well-defined and cheap to verify, keep your hands on it when the work is ambiguous or expensive to get wrong. An agent multiplies whatever clarity you bring. If the task is clear, it multiplies your speed. If the task is vague, it multiplies your confusion into a large, tidy-looking pull request.
The strongest version of the workflow pairs agentic coding with a written spec for anything non-trivial. A spec says what to build; the agent builds it; your review checks the build against the spec instead of against a vague memory. If you write software design documents or a technical spec already, you have most of the muscle. What changed is the reader: a spec used to inform a colleague who would ask when something was unclear. An agent does not ask. It guesses, confidently, so the spec has to close the gaps a human would have flagged.
Where agentic coding breaks
Every technique trades one set of risks for another. Agentic coding trades the slowness of typing for a set of quieter failures, and the quiet ones are the ones that cost you in a month instead of a minute.
| Failure mode | What it looks like | What to do |
|---|---|---|
| Review debt | The agent writes faster than you read, so you start skimming, then rubber-stamping. | Keep tasks small enough to review honestly. If you cannot review it, you cannot ship it. |
| Confident wrong guesses | The gaps you left in the task got filled with plausible defaults you never chose. | Frame the task: constraints, "done" criteria, out-of-scope. The gap is where the guess hides. |
| Hallucinated dependencies | The agent imports a package or calls an API method that does not exist but looks right. | Make "run it and show me the output" part of every task. Passing code cannot hallucinate. |
| Parallel systems | The agent rebuilds something you already have because it never saw the existing version. | Point it at the code to reuse in the task. An agent only knows the context you give it. |
| Rotting context | The AGENTS.md, specs, and docs the agent reads drift out of date, and the agent trusts them anyway. | Treat the agent's context as something to maintain, not just write. More on this next. |
The first three are about a single task and you catch them in review, if you actually review. The last two outlive the pull request, and they share a root cause: an agent is only as good as the context it reads, and that context both fills up with clutter and goes stale. The failure that gets the least attention is the one that compounds, so it gets its own section.
The context an agent reads is an artifact that rots
Here is the part almost nobody writing about agentic coding mentions. An agent does not know your codebase. It knows what it can read at the moment it runs: the files you point it at, the AGENTS.md in your repo, the spec attached to the task, the memory an AI keeps across sessions. Those are the agent's map of your world, and every one of them is a document that was true when someone wrote it.
The moment the code changes and the context does not, the map starts lying. Someone renames the Redis client, but AGENTS.md still says "see src/lib/redis.ts." A later PR changes the error format, but the spec the next agent reads still describes the old one. Nothing broke. No test failed. The agent just inherited a confident, authoritative, wrong description of the system and built on top of it, which is exactly how outdated documentation does its damage, one stale line at a time.
This is knowledge decay, and agentic coding is unusually exposed to it for two reasons. First, the context is precise, and precise things falsify fast: "we use pnpm" is trivially wrong the day the repo switches to bun. Second, the reader is an agent, which will not raise an eyebrow at a stale instruction the way a human teammate would. It reads the map, trusts it, and acts. The parts that rot first are the ones tied to something with its own clock:
- A named file or function. Any instruction that points at a path is wrong the refactor after it was written.
- An external fact. "The payments API returns
status: paid" carries an expiry date the day that API changes, and no context file has an expiry field. - An out-of-scope line that quietly became in scope. "We do not support webhooks yet" is true for exactly one release.
The workflow only pays off if the agent's context stays true, which means someone, or something, has to notice when it stops being true. The same discipline that keeps a note honest, checking the source instead of trusting the copy, is what keeps an agent's context honest past its first use. That is the premise Scribelet is built on: the notes, docs, and facts your AI works from are things to maintain, not just store. When the reference material an agent leans on can flag its own staleness, a drifting instruction gets caught before it ships the wrong thing instead of quietly steering the next ten tasks wrong.
Frame the work, then keep the context honest
Agentic coding moves you up a level, from writing code to deciding what code should be and checking that you got it. Done well, it is real leverage: frame a well-scoped task, delegate it to an agent that can plan and run and test, and review the result against criteria you set in advance. Done carelessly, it is a fast way to generate a lot of code nobody understands.
The two disciplines that separate the outcomes are small. Keep tasks small enough to review honestly, because the review is the part that makes the speed safe. And treat the agent's context, the AGENTS.md, the specs, the memory, the docs it reads, the way you would treat any note you rely on: as something that was true when you wrote it and needs a look before you trust it again. Set up your first desk and see what it looks like when the source of truth your AI works from is something that tells you when it has gone stale, instead of aging in silence while an agent builds on it.
Share this article