Virtual threads promise something Java teams have wanted for a decade: the scalability of reactive programming with plain, blocking, readable code. In Spring Boot it is literally one property. But "one property" hides real changes in how your service behaves under load, and the teams that flip the switch without understanding them meet exhausted connection pools, overwhelmed downstream services and confusing thread dumps. This guide explains how virtual threads work, how to enable them safely in Spring Boot, and the lessons we learned running them in production.

What virtual threads are

A classic Java thread is a thin wrapper around an operating system thread. OS threads are expensive — each reserves memory for its stack and context switches cost CPU — so servers use pools of a few hundred threads. When every thread is blocked waiting for a database or an HTTP call, new requests queue up even though the CPU is idle.

Virtual threads, final since Java 21, are lightweight threads managed by the JVM (JEP 444). When a virtual thread blocks on I/O, the JVM unmounts it from its carrier OS thread and runs another virtual thread there. You can have millions of them. Code stays synchronous: repository.findById() blocks the virtual thread, not the OS thread.

The key point: virtual threads improve throughput for I/O-bound workloads. They do not make individual requests faster, and they do nothing for CPU-bound work.

Enabling them in Spring Boot

Since Spring Boot 3.2, a single property switches the embedded web server, @Async execution, scheduling and several integrations to virtual threads (Spring Boot reference: virtual threads):

spring.threads.virtual.enabled=true
# Virtual threads are daemon threads; keep the JVM alive for schedulers and similar
spring.main.keep-alive=true

Two notes from the Spring Boot documentation:

  • Virtual threads require Java 21, but Java 24 or later is strongly recommended.
  • Once enabled, properties that configure thread pools — such as Tomcat's server.tomcat.threads.max — no longer have an effect, because virtual threads run on a JVM-wide scheduler.

Why Java 24+ matters: the pinning problem

In Java 21, a virtual thread that blocked inside a synchronized block or method stayed pinned to its carrier OS thread. With a small number of carriers (by default, one per CPU core), a few pinned threads could stall the whole application. Many popular libraries used synchronized internally, so this was a real production issue.

JDK 24 fixed it: virtual threads can now block inside synchronized without pinning their carrier (JEP 491). Pinning still occurs when native code calls back into Java and blocks, and JDK Flight Recorder records a jdk.VirtualThreadPinned event for those cases. If you adopt virtual threads, run on Java 25 LTS — see Java 25 LTS: what's new.

Pitfall 1: connection pools become the bottleneck

With a platform thread pool of 200, your service could never run more than 200 concurrent database queries. With virtual threads, 5,000 concurrent requests can all try to get a database connection at once. The connection pool — HikariCP defaults to 10 connections — now limits throughput, and requests wait on it.

This is actually healthy: the pool, not the thread pool, should be the concurrency limit for the database, and a database performs best with a modest number of connections (HikariCP: about pool sizing). But you must size it deliberately and set a sensible connectionTimeout, so requests fail fast instead of piling up.

Pitfall 2: you can now overwhelm downstream services

Thread pools were an accidental rate limiter. Remove them, and a traffic spike passes straight through to the payment provider, the ERP API or the legacy SOAP service that falls over at 50 requests per second. Add explicit limits:

@Component
class ErpClient {
    // At most 20 concurrent calls to the ERP, regardless of how many virtual threads we have
    private final Semaphore permits = new Semaphore(20);
    private final RestClient rest;

    ErpClient(RestClient.Builder builder) {
        this.rest = builder.baseUrl("https://erp.example.com").build();
    }

    Order fetchOrder(String id) throws InterruptedException {
        if (!permits.tryAcquire(2, TimeUnit.SECONDS)) {
            throw new ServiceUnavailableException("ERP busy, try again");
        }
        try {
            return rest.get().uri("/orders/{id}", id).retrieve().body(Order.class);
        } finally {
            permits.release();
        }
    }
}

Resilience4j bulkheads and rate limiters, or Spring Framework 7's @ConcurrencyLimit, do the same declaratively.

Pitfall 3: ThreadLocal assumptions

Code that caches expensive objects in ThreadLocal — formatters, buffers, connections — assumed a few hundred long-lived threads. With a new virtual thread per request, those caches are created and thrown away constantly and can increase memory use. Replace them with shared thread-safe objects or pools, and use scoped values for request context in new code (JEP 506).

Pitfall 4: CPU-bound work

Virtual threads don't add CPU. Image processing, PDF generation, heavy JSON transformations or cryptography still compete for the same cores. Keep CPU-heavy tasks on a bounded platform-thread executor so they don't starve request handling.

Monitoring virtual threads

  • Thread dumps: jcmd <pid> Thread.dump_to_file -format=json dump.json includes virtual threads, which classic jstack output does not show in a useful way (Oracle: virtual threads guide).
  • JFR events: jdk.VirtualThreadPinned and jdk.VirtualThreadSubmitFailed reveal scheduling problems.
  • Pool metrics: watch HikariCP pending connections and acquisition time; they are now your primary saturation signal.
  • Tracing: distributed traces show where requests actually wait; see OpenTelemetry observability.

Virtual threads vs reactive vs coroutines

Approach Code style Best for
Virtual threads Plain blocking code, Spring MVC, JDBC Most CRUD and integration services
Reactive (WebFlux, Reactor) Operators and streams Streaming, backpressure-heavy pipelines
Kotlin coroutines Sequential code with suspend Kotlin teams needing structured concurrency

For new Spring MVC services with JDBC, virtual threads are our default. We keep reactive stacks where backpressure and streaming are core requirements, and use coroutines in Kotlin services, as discussed in Spring Boot + Kotlin in 2026. How to structure such a service internally is covered in Modular monolith with Spring Modulith.

A safe rollout plan

  1. Upgrade to Java 24+ (ideally 25 LTS) and Spring Boot 3.2+.
  2. Load-test the service as is and record throughput, latency, pool metrics and memory.
  3. Enable spring.threads.virtual.enabled and spring.main.keep-alive on one instance.
  4. Size the connection pool deliberately; add limits for every downstream dependency.
  5. Repeat the load test, compare, and look for jdk.VirtualThreadPinned events.
  6. Roll out gradually, watching error rates of downstream services.

How to benchmark the switch honestly

Comparisons between platform and virtual threads are easy to get wrong. A method that gives results you can trust:

  1. Use realistic downstream latency. Virtual threads shine when requests wait on I/O; a benchmark against an in-memory stub that answers in microseconds will show little difference.
  2. Keep everything else identical: same JDK, heap, container limits, connection pool size and dataset.
  3. Ramp load gradually and record throughput, p50/p95/p99 latency and error rate at each step, not just the peak.
  4. Watch the bottleneck move. With virtual threads, saturation shifts from the thread pool to the connection pool or the downstream service — record those metrics too.
  5. Run long enough — at least 15–30 minutes per step — to see garbage collection and memory behavior, not just the warm-up phase.

In our tests the biggest gains appear in services that call several slow dependencies per request; CPU-bound services show essentially no difference.

FAQ

Do virtual threads make my API faster? They increase how many concurrent requests a service can handle when it mostly waits on I/O. Individual request latency stays the same or improves slightly under load.

Do I need to change my code? Usually not, beyond adding concurrency limits and replacing ThreadLocal caches. That's the main advantage over reactive programming.

Are virtual threads production-ready? Yes, especially on Java 24 and later, where the synchronized pinning issue is resolved.

Should we migrate WebFlux services to virtual threads? Only if the reactive code is a maintenance burden and the service doesn't rely on streaming or backpressure. Measure first.

Do virtual threads work with Kotlin coroutines? Yes. You can run coroutines on a dispatcher backed by virtual threads, though in most Spring MVC services one model is enough.

Sources

  1. OpenJDK. JEP 444: Virtual Threads and JEP 491: Synchronize Virtual Threads without Pinning.
  2. OpenJDK. JEP 506: Scoped Values.
  3. Spring Boot. Reference: virtual threads and Task execution and scheduling.
  4. Oracle. Virtual threads, Java SE 25 core libraries guide.
  5. HikariCP. About pool sizing.