Almost every engineering team now uses AI coding assistants in some form. In the Stack Overflow Developer Survey 2025, 84% of respondents said they use or plan to use AI tools in their development process. Yet the same survey shows that more developers distrust the accuracy of these tools than trust it. Managers hear "we are 2x faster" from one engineer and "it wastes my time" from another. Both can be true. This playbook explains what the evidence actually says, how to roll assistants out without creating a security or quality problem, and how to measure whether the investment pays off for your team.
What the research says (and what it doesn't)
The productivity debate is full of anecdotes. A few studies are worth knowing because they measure something concrete.
- Controlled task, greenfield code. In a randomized experiment, developers with GitHub Copilot implemented an HTTP server in JavaScript 55.8% faster than the control group (Peng et al., 2023). The task was small, well specified and started from an empty repository — the ideal case for an assistant.
- Experienced developers, mature codebases. METR ran a randomized trial with 16 experienced open-source maintainers working on real issues in repositories they knew well. With AI tools allowed, they took 19% longer — while believing they had been about 20% faster (METR, 2025; summary).
- Organisational level. The DORA 2024 report found that a 25% increase in AI adoption was associated with a 7.2% decrease in delivery stability and a 1.5% decrease in delivery throughput, even as individuals reported higher productivity and job satisfaction (DORA 2024). The 2025 edition frames AI as an amplifier: strong engineering practices get stronger, weak ones get exposed (DORA 2025).
The pattern is consistent: assistants shine on well-scoped tasks with fast feedback, and they struggle — or create hidden costs — when context is large, requirements are implicit, and review capacity is the bottleneck. Self-reported speed-ups are not reliable evidence. Plan your rollout around that fact.
Three categories of tools, three different risk profiles
"AI assistant" now covers very different products. Treat them separately in policy and budget.
| Category | Examples | What it does | Main risk |
|---|---|---|---|
| Inline completion | Copilot completions, Cursor Tab, JetBrains AI | Suggests the next lines as you type | Subtle bugs accepted without reading |
| Chat in the IDE | Copilot Chat, Cursor Chat, Claude in the IDE | Explains code, drafts functions, answers questions | Confidently wrong answers about your codebase |
| Agentic tools | Claude Code, Copilot coding agent, Codex, Cursor agent | Reads the repo, edits many files, runs commands and tests | Large diffs, shell access, secrets exposure |
Agentic tools deliver the biggest gains and carry the most risk, because they execute commands. We cover how to work with them productively in agentic coding best practices.
Step 1: Decide what problem you are solving
"Everyone gets a license" is not a strategy. Pick two or three concrete pain points and evaluate tools against them:
- Onboarding — new engineers asking the assistant to explain modules, data flows and conventions.
- Boilerplate and glue code — DTOs, mappers, API clients, migrations, configuration.
- Tests — characterization tests for legacy code, extra edge cases for existing suites (see AI-generated unit tests).
- Migrations — framework upgrades and API deprecations across many files (see AI for legacy modernization).
- Review load — a first-pass reviewer that catches obvious issues before a human looks (see AI code review).
Write these down. They become your evaluation criteria and later your metrics.
Step 2: Settle the legal and security baseline first
Before any code leaves a developer's laptop, answer these questions in writing:
- Data retention and training. Does the vendor train on your code? What is the retention period for prompts? Business and enterprise plans of the major vendors typically exclude customer data from training by default, but check the actual terms you sign — not the marketing page.
- Which repositories are allowed. Client code under NDA, regulated data, cryptographic material and anything with personal data may need a separate decision. Some clients forbid third-party AI processing outright.
- Secrets. Assistants read
.envfiles, shell history and config. Use deny-lists in tool settings, keep secrets in a manager rather than files, and run secret scanning on every push. - Command execution. For agentic tools, decide which commands run without confirmation. Read-only commands and tests are usually fine;
git push, deployment scripts and database clients are not. - Licensing of generated code. Enable duplicate-detection or code-referencing filters where the vendor offers them and keep normal license scanning in CI.
If you operate in the EU, AI coding tools are generally minimal-risk under the AI Act, but data protection rules still apply to any personal data inside your code or logs. Our EU AI Act guide and privacy guide for LLM apps cover the details.
Step 3: Run a time-boxed pilot with a control
A pilot should produce evidence, not enthusiasm. A design that works for teams of 10–100 engineers:
- Duration: 6–8 weeks. The first two weeks are a learning curve; do not judge results there.
- Participants: a mix of seniorities and codebases, including at least one legacy service. Volunteers only will bias results upward.
- Baseline: collect the same metrics for 4–6 weeks before the pilot, or keep a comparable group without the tool.
- Training: a 90-minute session on prompting, context files and reviewing AI output. Teams that skip training get noticeably less out of the tools.
- Weekly notes: each participant logs one task where the tool helped and one where it hurt. These qualitative notes are often more useful than the numbers.
Step 4: Measure outcomes, not keystrokes
Vendor dashboards report "acceptance rate" and "lines suggested". These are activity metrics; they say nothing about value. Use a balanced set:
| Dimension | Metric | Source |
|---|---|---|
| Throughput | PRs merged per engineer per week, lead time for changes | Git hosting, DORA metrics |
| Quality | Change failure rate, reverted PRs, bugs per release | Incident tracker, CI |
| Review load | Time to first review, PR size, review rounds | Git hosting |
| Maintainability | Duplicate code, test coverage trend | Static analysis |
| Developer experience | Survey: focus time, frustration, perceived usefulness | Quarterly survey |
| Cost | Licenses plus API spend per engineer | Billing |
Watch PR size closely. Assistants make it cheap to write code and expensive to review it; if average diff size doubles while review time per PR stays flat, someone is not reading carefully. The DORA findings on stability point exactly at this mechanism.
Step 5: Write a short, enforceable usage policy
Long policies are ignored. One page is enough:
# AI assistant policy (v1.2)
1. Approved tools: Claude Code, GitHub Copilot (business plan). Others need approval.
2. Allowed repos: all internal repos except `payments-core` and client repos marked `no-ai`.
3. You own every line you commit. "The AI wrote it" is not a review argument.
4. Never paste secrets, customer personal data or production dumps into prompts.
5. Agent tools: no auto-approval for push, deploy, DB or network commands.
6. PRs with significant AI-generated code must include tests and a short
description of what was verified manually.
7. New dependencies suggested by AI must be checked to exist and be maintained.
Rule 7 is not paranoia: LLMs regularly suggest packages that do not exist, and attackers register those names. See security of AI-generated code.
Step 6: Invest in context, not just licenses
The biggest difference between teams that get value from assistants and teams that don't is how much context they give the tool. Practical investments:
- Repository instruction files (
CLAUDE.md,AGENTS.md,.github/copilot-instructions.md) describing build commands, architecture, conventions and "never do" rules (Claude Code memory docs). - Fast, reliable test commands. An agent that can run
make testin 30 seconds verifies its own work; one that needs a 20-minute pipeline guesses. - Up-to-date architecture notes and ADRs. These help humans and models equally.
- Tool integrations via the Model Context Protocol for tickets, docs and observability, so the assistant can read the issue and the logs instead of relying on a copy-paste.
These investments are what context engineering means in a development setting.
Step 7: Adjust the review process
When code becomes cheap, review becomes the constraint. Adapt:
- Keep PRs small. Ask agents to split work into reviewable commits; reject 2,000-line "AI refactors" unless they are mechanical and covered by tests.
- Require evidence. A PR description should say how the change was verified: tests added, manual checks, screenshots.
- Use AI as a first-pass reviewer, not the only one. It catches typos, missing null checks and inconsistent naming; humans focus on design and business logic.
- Track reverts and hotfixes by origin for a few months to see whether AI-heavy changes behave differently in production.
Cost: what to budget
Seat licenses for completion and chat tools are predictable. Agentic tools that call models through an API can vary from a few dollars to hundreds per engineer per month depending on usage. Set per-user budgets and alerts from day one, review spend monthly, and apply the same levers we describe in LLM cost optimization — mostly model choice and context size.
As a rule of thumb, if a tool saves a mid-level engineer two hours a month, it has already paid for a typical seat. The hard part is not the license; it is making sure the saved time is not spent later on debugging and review.
Common failure modes
- Silent quality decline. More code, fewer tests, bigger PRs. Catch it with the metrics above.
- Skill atrophy among juniors. Juniors who accept suggestions they cannot explain do not learn. Pair them with seniors and ask them to explain AI-generated code during review.
- Shadow AI. If the approved tool is too restricted, engineers use personal accounts. Make the approved path the easiest one.
- Everything through the agent. Some tasks — delicate concurrency, security-critical code, performance tuning — benefit from slower human thinking. Let engineers choose.
FAQ
Which tool should we pick? Pilot two tools on the same tasks. Differences between teams and codebases are larger than differences in vendor benchmarks. Agentic CLI tools suit engineers comfortable in the terminal; IDE-integrated tools have a lower learning curve.
Will AI assistants let us hire fewer developers? The evidence so far supports faster work on well-defined tasks, not a replacement for engineering judgment. Most teams use the capacity to reduce backlog and technical debt rather than headcount.
How do we handle client code? Ask the client. Put AI usage into the contract or the statement of work, and keep a per-repository opt-out label.
Do we need our own models? Rarely, for coding. Self-hosted models make sense mainly when data cannot leave your infrastructure; see self-hosting LLMs for the trade-offs.
How long until we see results? Expect a learning curve of 2–4 weeks and stable data after about two months.
Sources
- Stack Overflow. Developer Survey 2025: AI.
- Peng et al. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot.
- METR (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.
- Google Cloud DORA. Accelerate State of DevOps Report 2024.
- Google Cloud DORA. State of AI-assisted Software Development 2025.
- Anthropic. Claude Code: manage memory (CLAUDE.md).
- Anthropic. Claude Code best practices for agentic coding.