For years, the default assumption was that AI features are built in Python and called from Java over HTTP. For many enterprise teams that meant another service, another deployment pipeline and another language to maintain. Spring AI changes that: it brings model access, retrieval, tool calling and observability into the Spring Boot programming model, with the auto-configuration, dependency injection and testing support Java teams already know. This guide shows how we build AI features with Spring AI in production, using Kotlin examples (the Java versions are nearly identical).

What Spring AI is

Spring AI is a Spring project that provides portable abstractions for AI providers and the building blocks of LLM applications (reference documentation):

  • ChatClient — a fluent API for prompts, system messages, options, streaming and structured output.
  • Model abstractions for chat, embeddings, images and audio across providers: Anthropic, OpenAI, Azure OpenAI, Amazon Bedrock, Google Vertex AI, Mistral, Ollama and others.
  • Advisors — interceptors around model calls for memory, retrieval, logging and guardrails.
  • VectorStore abstraction with implementations for PGvector, Qdrant, Weaviate, Milvus, Redis, Elasticsearch and more.
  • Tool calling via annotated methods.
  • Observability through Micrometer, plus evaluation helpers and MCP client/server support.

The value is not magic; it is consistency. Switching from one model provider to another, or from a hosted model to a self-hosted one via Ollama or vLLM, becomes a configuration change.

Setup

Use the Spring AI BOM and the starter for your provider and vector store:

// build.gradle.kts
dependencies {
    implementation(platform("org.springframework.ai:spring-ai-bom:1.0.0")) // use the current release
    implementation("org.springframework.ai:spring-ai-starter-model-anthropic")
    implementation("org.springframework.ai:spring-ai-starter-vector-store-pgvector")
    implementation("org.springframework.boot:spring-boot-starter-web")
    implementation("org.springframework.boot:spring-boot-starter-actuator")
}
# application.yml
spring:
  ai:
    anthropic:
      api-key: ${ANTHROPIC_API_KEY}
      chat:
        options:
          model: ${CHAT_MODEL}
          max-tokens: 1024
    vectorstore:
      pgvector:
        initialize-schema: false   # manage schema with Flyway/Liquibase instead
        dimensions: 1024
        index-type: HNSW
        distance-type: COSINE_DISTANCE

Keep the model name in configuration so you can switch models per environment without code changes; see choosing an LLM.

ChatClient basics

@RestController
@RequestMapping("/api/assistant")
class AssistantController(builder: ChatClient.Builder) {

    private val chat = builder
        .defaultSystem(
            """
            You are an internal assistant for ACME's sales team.
            Answer in the user's language. If you are not sure, say so.
            """.trimIndent(),
        )
        .build()

    @PostMapping("/ask")
    fun ask(@RequestBody req: AskRequest): String =
        chat.prompt()
            .user(req.question)
            .call()
            .content() ?: ""
}

ChatClient.Builder is auto-configured for the provider on the classpath. Build one client per use case with its own defaults — system prompt, advisors, tools — rather than one global client. Production prompts deserve the same rigour as code; see prompt engineering for developers.

Structured output into Kotlin data classes

Spring AI can map responses directly into typed objects, generating format instructions and a JSON schema from the target class:

data class LeadQualification(
    val companyName: String?,
    val budgetEur: Int?,
    val timeline: Timeline,
    val fitScore: Int,          // 1..5
    val reasoning: String,
)
enum class Timeline { NOW, THIS_QUARTER, LATER, UNKNOWN }

fun qualify(emailText: String): LeadQualification =
    chat.prompt()
        .user { it.text("Qualify this inbound lead email:\n<email>{email}</email>").param("email", emailText) }
        .call()
        .entity(LeadQualification::class.java)
        ?: error("Model returned no result")

Use nullable fields and an UNKNOWN enum value for information that may be missing, and validate business rules after mapping (Bean Validation works well). Where the provider supports native schema-constrained output, enable it. The general principles are in structured outputs from LLMs.

RAG with advisors and pgvector

Advisors wrap model calls with additional behaviour. Two of the most useful are chat memory and retrieval:

@Configuration
class AssistantConfig {
    @Bean
    fun supportChat(builder: ChatClient.Builder, vectorStore: VectorStore, chatMemory: ChatMemory): ChatClient =
        builder
            .defaultSystem(SUPPORT_SYSTEM_PROMPT)
            .defaultAdvisors(
                MessageChatMemoryAdvisor.builder(chatMemory).build(),
                QuestionAnswerAdvisor.builder(vectorStore)
                    .searchRequest(SearchRequest.builder().topK(6).similarityThreshold(0.5).build())
                    .build(),
            )
            .build()
}

@Service
class SupportAssistant(private val supportChat: ChatClient) {
    fun answer(userId: Long, conversationId: String, question: String): String =
        supportChat.prompt()
            .user(question)
            .advisors { it.param(ChatMemory.CONVERSATION_ID, "$userId:$conversationId") }
            .call()
            .content() ?: ""
}

Ingestion uses Spring AI's ETL-style DocumentReader, TextSplitter and VectorStore.add(). For production quality, go beyond the defaults: structure-aware chunking, metadata for filtering by tenant and access level, hybrid search, reranking. These are the decisions that matter most; see RAG for business, embeddings and chunking and vector databases compared. Using PGvector keeps vectors in the same PostgreSQL you already back up and monitor.

Access control: apply metadata filters based on the authenticated user in the SearchRequest (filter expression), never rely on the prompt to hide documents.

Tool calling

Annotate methods with @Tool and pass instances to the prompt:

class OrderTools(private val orders: OrderService, private val customerId: Long) {

    @Tool(description = "Get status, items and delivery estimate of one order of the current customer.")
    fun getOrderStatus(
        @ToolParam(description = "Order number, digits only, e.g. 482193") orderNumber: String,
    ): OrderStatusDto =
        orders.findForCustomer(customerId, orderNumber)
            ?.toDto()
            ?: OrderStatusDto.notFound(orderNumber)
}

fun answerWithTools(customerId: Long, question: String): String =
    chat.prompt()
        .user(question)
        .tools(OrderTools(orderService, customerId))
        .call()
        .content() ?: ""

The customer ID comes from the security context, not from the model. Return compact DTOs rather than JPA entities — lazy-loaded relations and 40-field entities waste tokens and leak data. Tool design principles are in designing tools for LLM agents, and the security model in AI agent security.

To expose your Spring services to external AI clients — Claude Desktop, IDE assistants, other agents — Spring AI also provides MCP server support; see MCP for business.

Streaming

For chat UIs, stream tokens as server-sent events:

@GetMapping("/stream", produces = [MediaType.TEXT_EVENT_STREAM_VALUE])
fun stream(@RequestParam q: String): Flux<String> =
    chat.prompt().user(q).stream().content()

This works in Spring MVC as well as WebFlux. Check that reverse proxies do not buffer the response. On the frontend, any SSE client works; for a React/Next.js frontend, see building an AI chat in Next.js.

Concurrency: blocking calls and virtual threads

LLM calls are slow I/O — seconds, not milliseconds. With classic thread-per-request servers, a burst of AI requests can exhaust the thread pool. On Java 21+ with Spring Boot 3.2+, enabling virtual threads (spring.threads.virtual.enabled=true) lets blocking call() code scale without rewriting it reactively. Watch out for pinning in synchronized blocks and for limits on downstream resources; our production experience is in virtual threads in Spring Boot.

Also add timeouts, retries with backoff for 429/5xx responses and fallbacks between providers — Spring Retry and the provider-level retry settings in Spring AI help; see LLM API reliability.

Observability

Spring AI emits Micrometer observations for ChatClient calls, model requests, advisors, tool calls and vector store operations, including token usage. With Spring Boot Actuator and an OpenTelemetry exporter, they appear as metrics and traces next to the rest of your application:

  • Track token usage per endpoint and per tenant for cost control; see LLM cost optimization.
  • Trace RAG requests end to end: retrieval latency, number of documents, model latency.
  • Be deliberate about logging prompt and completion content — it is often disabled by default for privacy; enable it only where policy allows.

The broader approach is in LLM observability and OpenTelemetry.

Testing

  • Unit tests: mock ChatModel or wrap ChatClient usage in your own interface for deterministic tests of surrounding logic, tools and permission checks.
  • Integration tests: Testcontainers for PostgreSQL with pgvector and, if you self-host, Ollama with a small model.
  • Evaluation: Spring AI includes evaluators such as relevancy and fact-checking that use a model as a judge. Treat them as part of an eval suite with a golden dataset, run in CI on prompt or model changes; see LLM evals.

Spring AI vs LangChain4j

LangChain4j is the other major JVM option. Both are solid:

Aspect Spring AI LangChain4j
Integration style Spring Boot auto-configuration, Spring idioms Framework-agnostic; integrations for Spring, Quarkus, Micronaut
High-level API ChatClient + advisors AI Services (annotated interfaces)
Observability Micrometer built in Listeners; integrations vary
Ecosystem Part of the Spring portfolio Large community, many integrations

For teams already on Spring Boot, Spring AI is the natural default. For Quarkus or framework-agnostic libraries, LangChain4j fits better.

FAQ

Is Spring AI production-ready? Yes, since its 1.0 GA release in 2025. As with any fast-moving library, pin versions, read release notes and keep integration tests.

Can we use Spring AI with self-hosted models? Yes, through the Ollama starter or the OpenAI starter pointed at an OpenAI-compatible server such as vLLM.

Does it work with Kotlin coroutines? The blocking API works fine on virtual threads; the streaming API returns Reactor Flux, which converts to Kotlin Flow with asFlow().

Should AI logic live in a separate microservice? Not by default. A module inside your existing service — for example in a Spring Modulith architecture — is simpler until scaling or team boundaries justify a split.

Sources

  1. Spring. Spring AI project and reference documentation.
  2. Spring AI. ChatClient API.
  3. Spring AI. Tool calling.
  4. Spring AI. Vector databases.
  5. LangChain4j documentation.
  6. pgvector.