Coding agents are different from autocomplete. Tools such as Claude Code, OpenAI Codex, the Copilot coding agent or Cursor's agent mode read your repository, edit dozens of files, run commands, look at the output and try again. Used well, they turn a half-day chore into a twenty-minute review. Used badly, they produce sprawling diffs that look plausible and break things in ways nobody notices until production. This guide collects the practices that, in our experience and in the published guidance from tool vendors, separate the two outcomes.
How a coding agent actually works
Under the hood, every coding agent is the same loop: the model receives a task plus context, decides on an action (read a file, search, edit, run a command), the harness executes it, and the result goes back into the context. The loop continues until the model decides it is done or hits a limit. This is the "agent" pattern described in Anthropic's Building effective agents and formalised in research such as ReAct.
Two consequences follow:
- The agent only knows what is in its context window. It does not "remember" your conventions unless they are written down somewhere it reads.
- The agent is only as good as its feedback. If it can run tests, a type checker and a linter, it can correct its own mistakes. If it cannot, it guesses.
Research on SWE-agent showed that the design of the agent-computer interface — how files are shown, how edits are applied, what feedback commands return — strongly affects success rates on real GitHub issues (Yang et al., 2024). Most of the practices below are about improving that interface for your repository.
Practice 1: Write a project instruction file
Every serious agent supports a repository-level instruction file that is loaded into context at the start of each session: CLAUDE.md for Claude Code, AGENTS.md for Codex and several other tools, .github/copilot-instructions.md for Copilot. Keep it short, specific and maintained.
# CLAUDE.md
## Commands
- Install: `pnpm install`
- Dev server: `pnpm dev` (port 3000)
- Tests: `pnpm test` (Vitest, ~20 s). Single file: `pnpm test path/to/file.test.ts`
- Types: `pnpm tsc --noEmit` — must pass before you say you're done
- Lint: `pnpm lint --fix`
## Architecture
- Next.js App Router in `app/`, shared UI in `components/`, domain logic in `lib/`
- Server-only code imports "server-only"; never import it from client components
- All money values are integers in minor units (cents)
## Rules
- Do not edit generated files in `lib/api/generated/`
- Do not add dependencies without asking
- Prefer small commits with conventional commit messages
What to include: commands, architecture in a few lines, non-obvious conventions and hard "never" rules. What not to include: generic advice the model already follows ("write clean code"), long style guides that the linter enforces anyway, or secrets. Anthropic's Claude Code best practices recommends iterating on this file like a prompt: when the agent repeats a mistake, add one line that prevents it.
Practice 2: Explore, plan, then execute
The most common failure is letting the agent start editing immediately. For anything bigger than a one-file fix, split the session into phases:
- Explore. "Read the payment module and the tests for refunds. Don't write code yet. Summarise how refunds flow through the system."
- Plan. "Propose a plan to support partial refunds. List the files you will change and the tests you will add." Review the plan. Correct misunderstandings here — it is ten times cheaper than correcting code.
- Execute. "Implement step 1 of the plan. Run the tests after each step."
- Verify and commit. "Run the full test suite, type check and lint. Summarise what changed and what you verified."
Many tools have a dedicated plan mode that prevents edits until you approve. Use it. For complex tasks, ask the agent to write the plan to a file (docs/plans/partial-refunds.md); it survives context resets and doubles as documentation.
Practice 3: Give the agent a way to verify its work
Agents perform dramatically better when they can check results themselves. Concretely:
- Fast tests. Make sure a targeted test run takes seconds. If the full suite is slow, document how to run a single file or test.
- Test-first for new behaviour. Ask the agent to write failing tests, confirm they fail, commit them, then implement until they pass. This prevents the classic failure where the agent "fixes" the test to match buggy code.
- Type checker and linter as part of the definition of done.
- Visual feedback for UI. For frontend work, a browser automation tool or screenshot lets the agent compare its result with a design.
- Reproduction scripts for bugs. "Write a script that reproduces the bug, confirm it fails, then fix it."
We go deeper into test generation in AI-generated unit tests.
Practice 4: Manage context deliberately
Long sessions degrade. The context fills with stale file contents, failed attempts and tool output, and the model's attention to early instructions drops. Practical habits:
- One task per session. Clear the context between unrelated tasks.
- Point to files explicitly rather than asking the agent to "look around".
- Use subagents for research. Many tools can spawn a subagent with a fresh context to investigate a question and return only the conclusion (Claude Code subagents). This keeps the main context focused.
- Compact or summarise when a session gets long, and restart from the plan file if quality drops.
The underlying principles are covered in context engineering for AI agents.
Practice 5: Set permissions you would set for a new contractor
An agent with shell access can do anything your user account can do. Configure permissions explicitly:
| Action | Recommended default |
|---|---|
| Read files in the repo | Allow |
| Edit files in the repo | Allow, review the diff |
| Run tests, linters, type checker | Allow |
| Install dependencies | Ask |
git commit |
Allow on feature branches |
git push, open PR |
Ask |
Network calls, curl, cloud CLIs |
Ask or deny |
| Database clients, deploy scripts | Deny |
Read .env, ~/.ssh, credential files |
Deny |
Run agents in a container or a dev VM for unattended work, especially "YOLO" modes that skip confirmations. Remember that the agent may read untrusted content — issue text, web pages, dependency READMEs — that contains instructions. That is prompt injection, and the defences are architectural; see AI agent security.
Hooks are another lever: deterministic scripts that run before or after tool calls, for example to run the formatter after every edit or block writes to protected paths (Claude Code hooks). Hooks enforce rules the model might forget.
Practice 6: Keep diffs reviewable
The agent writes code faster than you can review it. Protect your review capacity:
- Scope tasks narrowly. "Add partial refunds to the API" is better than "improve the payment module".
- Ask for small, logically separated commits. Mechanical renames in one commit, behaviour changes in another.
- Ban drive-by refactors in the instruction file: "Do not refactor code unrelated to the task."
- Review like a human wrote it. Read every line of non-mechanical changes. Check error handling, edge cases and whether the tests actually test the behaviour.
- Ask the agent to review its own diff with a fresh context before you do. It catches a surprising number of issues; see AI code review.
Practice 7: Know where agents struggle
Agents are excellent at well-specified, verifiable work: migrations, adding endpoints that follow existing patterns, writing tests, fixing failing builds, explaining unfamiliar code. They struggle with:
- Ambiguous requirements. They will pick an interpretation and commit to it confidently.
- Cross-cutting design decisions. Choosing an architecture needs trade-offs only you know.
- Concurrency and performance problems that need measurement, not reasoning.
- Huge files and generated code that exceed what fits comfortably in context.
- Tasks without feedback, such as code that only runs in production.
For these, use the agent as a research assistant — "find all places where we lock this table" — and keep the decisions human.
Practice 8: Run agents in parallel, carefully
Once a team trusts its workflow, parallel agents become attractive: several sessions working on independent tasks in separate Git worktrees or cloud sandboxes. This works when tasks are truly independent and each has its own verification. It fails when agents touch the same files or when reviews pile up. Start with two parallel sessions, not ten. For background agents triggered from issues or CI, see AI agents in CI/CD.
A session template
Task: Support partial refunds for card payments (issue #482).
Context:
- Refund logic: lib/payments/refunds.ts, tests in lib/payments/__tests__/
- Provider API docs: docs/providers/stripe-refunds.md
- Constraint: amounts are integers in minor units
Process:
1. Read the files above and summarise the current refund flow. Stop.
2. After my approval, write failing tests for partial refunds. Stop.
3. Implement until tests pass. Run `pnpm test lib/payments` and `pnpm tsc --noEmit`.
4. Summarise changes and anything you were unsure about.
FAQ
Should juniors use agents? Yes, with guardrails: they must be able to explain every change in review. Agents are great tutors when asked "why", and poor ones when used only to produce output.
How do we stop the agent from changing tests to make them pass? Commit tests first, state in the instruction file that tests are the specification, and review test diffs separately.
Which model should the agent use? The most capable model for planning and complex changes; a faster, cheaper one for mechanical edits if the tool lets you switch. See choosing an LLM.
Is it safe to let an agent work unattended? Only in an isolated environment without production credentials, on a branch, with the result going through normal review.
Sources
- Anthropic. Claude Code best practices for agentic coding.
- Anthropic. Building effective agents.
- Yang et al. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.
- Yao et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models.
- Anthropic. Claude Code documentation: subagents and hooks.
- AGENTS.md — an open format for guiding coding agents.