Test generation looks like the perfect job for an AI assistant: tedious, pattern-based and easy to verify. Ask any modern model to "write unit tests for this class" and you will get a tidy file with a dozen green tests in seconds. The problem is that green tests are not the goal. Tests exist to fail when the code is wrong, and AI-generated tests often do not. This guide explains why, and how to set up a workflow where generated tests actually raise your defect detection rather than just your coverage number.
The core problem: tests that mirror the implementation
When a model generates tests by reading the implementation, it tends to encode what the code does, not what it should do. If the function has a bug, the generated test asserts the buggy behaviour. Coverage rises, confidence rises, and the bug is now protected by a test.
Typical symptoms of low-value generated tests:
- Assertions that restate the implementation (
expect(calc(2)).toBe(2 * RATE)copied from the code). - Heavy mocking, so the test only checks that mocks were called in order.
- Only happy paths; no boundaries, nulls, empty collections or error cases.
- Snapshot tests of large objects that nobody reads.
- Tests that pass even when you delete the body of the method.
The last symptom is measurable, and that measurement is the key to the whole workflow.
Lessons from Meta's TestGen-LLM
Meta published one of the most informative industrial reports on LLM test generation. Their TestGen-LLM tool generated additional tests for existing Kotlin test classes in Instagram and Facebook, but it did not trust the model. Every candidate had to pass a series of filters: it must build, it must pass reliably across several runs (to remove flaky tests), and it must increase coverage. Only then was it proposed to engineers as a diff.
The numbers are instructive. Of generated test cases, about 75% built correctly, 57% passed reliably, and 25% increased coverage. In Meta's test-a-thons, engineers accepted 73% of the recommended improvements for production (Alshahwan et al., 2024).
The lesson: most raw generations are not useful, and that is fine, as long as automatic filters throw them away. Treat the model as a candidate generator and put deterministic checks between it and your repository.
Mutation testing: the quality gate that matters
Coverage tells you which lines ran. Mutation testing tells you whether your tests would notice if those lines were wrong. A mutation tool makes small changes to your code — flipping > to >=, replacing a return value, removing a call — and runs the tests. If the tests still pass, the mutant "survived", and you have found a gap.
Mature tools exist for most ecosystems:
| Language | Tool |
|---|---|
| Java, Kotlin (JVM) | PIT (pitest) |
| JavaScript, TypeScript, C#, Scala | Stryker |
| Python | mutmut, Cosmic Ray |
| Go | go-mutesting, Gremlins |
Mutation score is a much better target for AI-generated tests than coverage, because it is hard to game: a test that asserts nothing kills no mutants. Research on combining LLMs with mutation testing — feeding surviving mutants back to the model and asking for a test that kills them — reports clear improvements in fault detection over coverage-guided generation, and this is now a practical workflow you can run with any agent.
A workflow that works
Step 1: Write the specification before the tests
Do not ask the model to "test this class". Give it the intended behaviour:
Write JUnit 5 tests for DiscountService.calculate(order).
Business rules (source of truth — do NOT infer rules from the implementation):
- Orders >= 1000 UAH get 5% off; >= 5000 UAH get 10% off.
- Loyalty members get an extra 2%, applied after the volume discount.
- Discounts never exceed 15% in total.
- Orders with any item on promotion get no volume discount.
- Amounts are BigDecimal, rounded HALF_UP to 2 decimals.
Include boundary values (999.99, 1000.00, 4999.99, 5000.00),
an empty order and a null customer. Use AssertJ. No mocks for value objects.
When the spec and the implementation disagree, the test fails — which is exactly what you want. Either the code has a bug or the spec is outdated, and a human decides.
Step 2: Run, filter, and keep only valid candidates
Automate the TestGen-LLM filters:
- Compiles — discard otherwise, or feed the compiler error back once.
- Passes reliably — run three to five times; discard flaky tests.
- Fails on the right thing — for each failing test, a human checks whether the code or the test is wrong.
Step 3: Run mutation testing and loop on survivors
# JVM: run PIT for one package
./gradlew pitest -Ppitest.targetClasses='com.shop.discount.*'
# TypeScript: run Stryker for one directory
npx stryker run --mutate "src/discount/**/*.ts"
Feed surviving mutants back to the agent: "This mutant survived: in DiscountService.java:42, >= was changed to >. Write a test that fails on the mutant and passes on the original." Repeat until the mutation score stops improving or reaches your threshold.
Step 4: Review like production code
Generated tests become part of the codebase and must be maintained. Check readability, names that describe behaviour (appliesTenPercentAtExactlyFiveThousand), absence of over-mocking and that each test has one reason to fail.
Characterization tests for legacy code
For legacy code without a spec, the "mirror the implementation" problem becomes a feature. Characterization tests, a term from Michael Feathers, deliberately capture current behaviour — bugs included — so you can refactor safely. LLMs are excellent at this:
- Ask the agent to enumerate input classes from the code's branches.
- Generate tests that record current outputs, including surprising ones.
- Mark suspicious results with a comment like
// CHARACTERIZATION: looks like a bug, see #123. - Use mutation testing to confirm the suite actually pins the behaviour.
This is the foundation for any migration or modernization; see AI for legacy code modernization.
Prompts and patterns by stack
Java/Kotlin with Spring. Ask for plain unit tests of domain classes first; use @WebMvcTest or @DataJpaTest slices only where needed. Explicitly forbid @SpringBootTest for unit tests — models love it, and it makes suites slow. For Kotlin, specify the framework (JUnit 5 + AssertJ, or Kotest) and MockK instead of Mockito. Our Spring Boot with Kotlin guide covers the setup.
TypeScript. Specify Vitest or Jest, testing-library for React components and "test behaviour visible to the user, not implementation details". Ask for test.each tables for boundary values — models produce readable tables and they cover many cases cheaply.
Property-based tests. Ask the model to propose invariants ("total never negative", "discount never above 15%") and implement them with jqwik, Kotest property testing, fast-check or Hypothesis. LLMs are good at finding invariants humans forget.
Agentic test generation at scale
With a coding agent and a fast test command, you can run test generation as a background task over a whole module: the agent picks the least-covered class, writes tests against a spec or characterizes behaviour, runs the filters and mutation tests, and opens a small PR per class. Keep the PRs small and reviewable, as described in agentic coding best practices, and schedule it outside business hours if your CI capacity is limited. For running agents in pipelines, see AI agents in CI/CD.
Wiring the quality gate into CI
Generated tests should pass through the same gates as the generator's filters, every time they change. A practical setup:
# .github/workflows/test-quality.yml (excerpt)
- name: Unit tests, three runs to detect flakiness
run: for i in 1 2 3; do ./gradlew test --rerun-tasks || exit 1; done
- name: Mutation testing on changed modules
run: ./gradlew pitest -Ppitest.mutationThreshold=70
Set the mutation threshold per module rather than globally: legacy modules may start at 40% and rise over time, while new domain code can hold 75–80%. Run full mutation testing nightly if it is too slow for every pull request, and fail the build only on decreases, not on modules that have not reached the target yet. This turns mutation score into a ratchet: generated tests can only raise it.
Metrics to track
- Mutation score for modules with AI-generated tests, before and after.
- Escaped defects — bugs found in production in areas with generated tests.
- Flaky test rate — generated tests must not make CI less reliable.
- Test suite runtime — watch for slow integration tests disguised as unit tests.
- Review time per test PR — if reviewers spend longer than writing tests by hand would take, adjust the workflow.
Evaluating the test generator itself is a small eval problem; the approach in LLM evals applies directly.
FAQ
Should AI generate tests for code it just wrote? It can, but the tests will share the code's misunderstandings. Write or review the spec yourself, or generate tests in a fresh session from the spec only, without showing the implementation.
Is 100% coverage a good goal for generated tests? No. It encourages tests of trivial getters and generated code. Aim for high mutation scores in business logic.
How do we avoid brittle tests? Forbid mocking of value objects and internal collaborators, prefer testing through public APIs, and reject snapshot tests unless the output is small and meaningful.
Does it work for integration tests? Yes, especially with Testcontainers and clear fixtures. Give the agent the docker-compose setup and a working example test to imitate.
Sources
- Alshahwan et al. (2024). Automated Unit Test Improvement using Large Language Models at Meta.
- PIT mutation testing for the JVM.
- Stryker Mutator.
- Petrović, G., Ivanković, M. (2018). State of Mutation Testing at Google. ICSE-SEIP.
- Michael Feathers. Working Effectively with Legacy Code. Prentice Hall, 2004.
- Anthropic. Claude Code best practices: write tests, commit; code, iterate, commit.