Every company with software older than five years has a modernization backlog: a Java 8 service nobody dares to touch, a framework two major versions behind, a test suite on a deprecated library, a monolith that should have been split years ago. These projects get postponed because they are expensive, risky and boring. AI coding agents change the economics of the boring part. They do not remove the risk. This guide shows where AI helps in modernization, what large companies have published about doing it at scale, and a step-by-step process that keeps the risk under control.
What the industry has learned
Several organisations have published concrete results from AI-assisted migrations:
- Google used LLMs for large internal code migrations, such as changing integer types for IDs across a huge codebase. In their report, roughly 80% of the code modifications in landed changes were AI-authored, and engineers estimated the total migration time was reduced by about 50% compared with doing it manually (Nikolov et al., 2025). Crucially, the pipeline combined LLM edits with deterministic tooling to find locations, builds and tests to validate, and human review.
- Airbnb migrated thousands of React test files from Enzyme to React Testing Library using an LLM-driven pipeline with automated validation and retry steps, completing in weeks what had been estimated at more than a year of manual work (Airbnb Engineering, 2025).
- AWS offers agent-driven transformations for Java version upgrades and .NET porting in Amazon Q Developer, again combining models with build and test verification (AWS documentation).
The common pattern across all of them: the LLM is one stage in a pipeline. Deterministic tools find what to change, the model makes context-sensitive edits, and builds and tests decide whether a change is accepted. Nobody reported success from "ask a chatbot to rewrite the service".
Types of modernization and how much AI helps
| Type | Example | AI contribution |
|---|---|---|
| Runtime upgrade | Java 8 → 21/25, Node 16 → 22, Python 3.8 → 3.12 | High: mechanical edits plus fixing compile and test failures |
| Framework upgrade | Spring Boot 2 → 3, Angular → newer, Odoo 16 → 18 | High, combined with recipes and release notes |
| Library replacement | Enzyme → RTL, Joda-Time → java.time, Moment → date-fns | Very high: well-defined target patterns |
| Language migration | Java → Kotlin, JavaScript → TypeScript | High for syntax, medium for idioms |
| Architecture change | Monolith → modules or services | Medium: analysis and boilerplate, humans own design |
| Platform migration | 1C → Odoo, custom ERP → standard | Medium: data mapping and logic extraction |
| Full rewrite | COBOL → Java | Low to medium: behaviour discovery is the hard part |
The further down the table, the more the work is about understanding behaviour rather than transforming syntax, and the more human judgement it needs.
Step 1: Inventory and understand before changing anything
Start with analysis. Agents are excellent research assistants on unfamiliar code:
- Dependency and API inventory. Which deprecated APIs are used, where, and how often? Combine
grep, build tools (mvn dependency:tree,gradle dependencies) and the agent's summaries. - Behaviour documentation. Ask the agent to document each module: inputs, outputs, side effects, external calls, implicit business rules. Review these documents with someone who knows the system; they will find misunderstandings early.
- Risk map. Which areas have no tests, high change frequency or many incidents? Those need the most care.
Store the results in the repository (docs/modernization/). They become context for every later agent session and for new team members.
Step 2: Build the safety net — characterization tests
You cannot safely change code whose behaviour you cannot verify. Before any transformation, generate characterization tests that pin the current behaviour of the areas you will touch. This is one of the most valuable uses of AI in modernization; the method is described in detail in AI-generated unit tests. Use mutation testing to confirm the tests actually detect changes.
For services with external behaviour that matters more than internal structure, add golden master tests at the API level: record requests and responses from the current system and replay them against the new one.
Step 3: Use deterministic tools for the mechanical part
Do not ask an LLM to do what a refactoring tool does reliably. For JVM projects, OpenRewrite provides tested recipes for upgrades such as Java versions, Spring Boot 3, Jakarta EE namespaces and JUnit 5:
# Example: apply OpenRewrite's Spring Boot 3 upgrade recipe with Maven
mvn -U org.openrewrite.maven:rewrite-maven-plugin:run \
-Drewrite.recipeArtifactCoordinates=org.openrewrite.recipe:rewrite-spring:RELEASE \
-Drewrite.activeRecipes=org.openrewrite.java.spring.boot3.UpgradeSpringBoot_3_0
Similar tools exist elsewhere: codemods with jscodeshift or ts-morph for JavaScript, pyupgrade and LibCST for Python, the .NET Upgrade Assistant. Run them first; let the agent handle what is left — the compile errors, the test failures and the places where context matters.
Step 4: Let the agent work in small, verified batches
The transformation loop that works:
- Pick a small unit: one module, one package or 10–20 files.
- Give the agent the target pattern with examples, the release notes, and the commands to build and test.
- The agent applies changes and iterates until build and tests pass.
- A human reviews the diff, with focus on behaviour changes rather than syntax.
- Merge and deploy behind a flag or to a canary if possible.
- Repeat.
Task: migrate package com.shop.billing from Joda-Time to java.time.
Rules:
- Follow docs/modernization/joda-to-javatime.md (mapping table + examples).
- Preserve time zone semantics exactly; billing uses Europe/Kyiv.
- Do not change public method signatures in this batch; add adapters if needed.
- After changes run: ./gradlew :billing:test :billing:pitest
Stop and ask if a conversion is ambiguous (e.g. LocalDate vs ZonedDateTime).
The "stop and ask" rule matters. Ambiguity is where silent bugs come from, and agents will resolve it confidently unless told otherwise. Our general workflow for agents is in agentic coding best practices.
Step 5: Roll out with the strangler fig, not a big bang
For architectural changes and platform migrations, the strangler fig pattern remains the safest approach: route a slice of functionality to the new implementation, verify in production, and expand. AI makes it cheaper to build each slice, which makes incremental migration more attractive than ever.
Useful techniques:
- Parallel run. Execute old and new code paths for the same request, compare results, return the old one. Log differences and investigate them.
- Feature flags per migrated slice, with fast rollback.
- Data migration rehearsals on production-size copies before cutover.
For Java services, moving toward a modular monolith with Spring Modulith is often a better first target than microservices — and agents are good at enforcing module boundaries once they are defined.
Example: a Java 8 to Java 25 upgrade
A typical sequence we use for JVM upgrades:
- Upgrade build tooling (Maven/Gradle wrapper, plugins) and get the build green on the old JDK.
- Run OpenRewrite recipes for the target Java version and for Jakarta EE if moving to Spring Boot 3.
- Switch the JDK in CI; let the agent fix compile errors in batches.
- Fix test failures — often reflection, removed APIs, changed default behaviour (for example, locale data, TLS defaults).
- Adopt new features selectively: records for DTOs, pattern matching, and, where it fits, virtual threads.
- Load test before production; GC and memory behaviour change between versions.
The details of the target release are covered in Java 25 LTS: what's new.
Platform migrations: where AI helps less than you expect
Migrating business systems — for example from 1C to Odoo — is mostly about data mapping, business process decisions and user training. AI helps with:
- Extracting business rules from old code and reports into readable documents.
- Generating data mapping scripts and validation queries.
- Drafting custom modules that follow the target platform's conventions.
It does not replace workshops with users or the decision about which customisations to drop.
Measuring progress
- Migration completion by module or by count of deprecated API usages remaining.
- Change failure rate of migration PRs versus regular PRs.
- Behavioural differences found in parallel runs.
- Engineer hours per unit — compare early batches with later ones; well-tuned pipelines get much faster.
FAQ
Can we just ask an agent to rewrite the whole application in a new stack? You will get code that looks complete and behaves differently in hundreds of small ways. Rewrites need behaviour discovery and incremental verification; AI speeds up both but does not skip them.
Which model should we use? Use a strong model for analysis and ambiguous transformations, and consider a cheaper one for high-volume mechanical batches once the pattern is proven. See choosing an LLM.
How do we handle code with no tests at all? Characterization tests first. If even that is impractical, start with API-level golden master tests and parallel runs.
Is the generated code maintainable? It follows the examples you give. Invest in a good mapping document with idiomatic before/after samples, and the output will match.
Sources
- Nikolov et al. (2025). How is Google using AI for internal code migrations?
- Airbnb Engineering (2025). Accelerating Large-Scale Test Migration with LLMs.
- AWS. Amazon Q Developer: transforming code.
- OpenRewrite documentation.
- Martin Fowler. Strangler Fig Application.
- Michael Feathers. Working Effectively with Legacy Code. Prentice Hall, 2004.