What Spring AI Does Between chatClient.prompt() and the Answer (Memory, RAG, Tool Loop)

October 1, 2026 · How It Actually Works: Real Systems as State Machines (part 9)

▶ Watch on YouTube & subscribe to The Stack Underflow

Your controller calls chatClient.prompt(), and a few seconds later an answer comes back. In between, Spring AI loads the conversation’s memory, searches your documents, lets the model ask for tools, runs them in your code, and saves the result. The video follows one request through all of it as a hierarchical state machine. This page is the written companion: the same path, the defaults that decide it, and the places it can fail.

The one-line version: advisor order decides what runs once and what runs on every model call: memory sits outside the tool loop, retrieval sits inside it, and the model only asks for tools that your code runs.

Last verified against the Spring AI 2.0.1 reference documentation and the v2.0.1 source: 1 October 2026. Scope, as in the video: Spring AI 2.0.1 on Spring Boot 4, a blocking ChatClient call from a Spring MVC controller, with MessageChatMemoryAdvisor, RetrievalAugmentationAdvisor, the auto-registered ToolCallingAdvisor, one local @Tool method and one MCP tool. Not covered: streaming, calling ChatModel directly, Spring AI 1.x.

Six words, six memory pegs

The video opens with a recap from the Spring AI words primer. The pegs are analogies, memory aids rather than facts:

WordMemory pegWhat it actually is
Modelthe consultant with no memoryRuns at the provider; stateless between calls
ChatClient (ChatModel)the assistant (the phone line)The fluent API your service calls; ChatModel is the lower-level interface it uses
Advisorthe checklistIntercepts the request and response; the chain is a stack
Chat memorythe notes folderRecent messages put back into each prompt
RAGcards clipped to the briefRetrieved document chunks added to this prompt
Toolyour staff do the lookupThe model asks; your code runs it

Startup: Spring AI is just beans

Spring AI starts no runtime of its own. The starters you add (a model provider, a vector store, the MCP client) contribute auto-configuration, and Spring Boot creates ordinary beans: a ChatModel and an EmbeddingModel for your provider, a VectorStore, a ChatMemory, and a prototype ChatClient.Builder your service uses to build its ChatClient.

Each configured MCP connection gets a client, over STDIO, SSE or Streamable HTTP. By default (spring.ai.mcp.client.initialized=true) it is initialized when it is created: in the MCP spec revision this client speaks (2025-11-25), the initialize request establishes the protocol version and exchanges capabilities. The client then lists the server’s tools, and with spring.ai.mcp.client.toolcallback.enabled=true (the default) Spring AI exposes them through a ToolCallbackProvider bean, so remote tools look like local ones.

Separately, an ingestion pipeline fills the vector store: a DocumentReader such as PagePdfDocumentReader reads a source, a TokenTextSplitter cuts it into chunks, and the VectorStore (which is also a DocumentWriter) embeds each chunk and stores it with its text and metadata. Ingestion runs when you decide, not inside a chat request. If it fails, chat requests aren’t affected; they just can’t find those documents.

One request: call() sends nothing

Your service fills the briefing pack: a system message, the user’s message, and the conversation ID as an advisor parameter. Then:

OrderAnswer answer = chatClient.prompt()
    .system(SUPPORT_RULES)
    .user(question)
    .advisors(a -> a.param(ChatMemory.CONVERSATION_ID, conversationId))
    .call()                       // nothing sent yet
    .entity(OrderAnswer.class);   // the advisor chain and the model call run here

The ChatClient reference is explicit: “Calling the call() method does not actually trigger the AI model execution. Instead, it only instructs Spring AI whether to use synchronous or streaming calls.” The request happens when you ask for the result: content(), entity(), chatResponse() or responseEntity().

The advisor chain is a stack

Advisors run like a stack: the one with the lowest order value sees the request first and the response last. With the defaults, three advisors matter, and their order decides everything that follows:

AdvisorDefault orderWhere it runsHow often per request
MessageChatMemoryAdvisorHIGHEST_PRECEDENCE + 200outside the tool looponce
ToolCallingAdvisor (auto-registered)HIGHEST_PRECEDENCE + 300it is the loop—
RetrievalAugmentationAdvisor0inside the tool looponce per model call

HIGHEST_PRECEDENCE is Integer.MIN_VALUE, so an order of 0 is far larger than HIGHEST_PRECEDENCE + 300, which places the retrieval advisor inside the loop. The tool-calling advisor page states the memory case directly: its default order “is higher than the default MessageChatMemoryAdvisor order (HIGHEST_PRECEDENCE + 200), which places memory advisors outside the tool loop by default.”

Memory: once, outside the loop

The memory advisor needs a conversation ID on every call, the case number for whose notes folder to load. In Spring AI 2.0.1 there is no default: a call without one throws IllegalArgumentException. With an ID, it reads the conversation’s messages from ChatMemory (by default a MessageWindowChatMemory keeping the last 20) and puts them into the prompt.

It saves the user message to memory on the way in, right after loading the history, and the final answer on the way out (before() and after() in the v2.0.1 source). The intermediate tool-call messages are not stored in chat memory; the docs list this as a current limitation. ToolCallingAdvisor keeps them only for the duration of its loop.

The tool loop, with retrieval inside it

Each pass of the loop goes: retrieval, then the model, then (maybe) tools.

Retrieval (RAG). Optional query transformers can rewrite the question first. The VectorStoreDocumentRetriever embeds the query and asks the vector store for the closest chunks, limited by top K, a similarity threshold and a metadata filter. That filter is where you enforce which documents this user may see. The ContextualQueryAugmenter adds the documents to the user message. If nothing is found, by default it tells the model not to answer rather than guess; allowEmptyContext(true) turns that off. RAG does not train the model; it only adds text to this prompt.

The model call. The ChatModel turns the prompt into the provider’s request, including the tool definitions. In Spring AI 2 the ChatModel does not execute tools: it returns the model’s response. If the response has no tool calls, it is the final answer and the loop ends.

If the provider fails, the exception reaches your code. Retries are provider-specific. Spring AI’s own retry handling (in the v2.0.1 source) maps 5xx errors to a retryable TransientAiException and 4xx errors to NonTransientAiException unless you configure on-client-errors or on-http-codes, so by default not even a 429 rate limit is retried.

Running a tool. The ToolCallingManager looks up the tool callback by name:

  • A local @Tool method runs in your app, with the arguments the model chose plus a ToolContext for data such as the tenant or user. ToolContext is never sent to the model. Authorization belongs here, in your code, not in the prompt.
  • An MCP tool goes through the MCP client as a tools/call request; the server runs it in its own process.
SituationDefault behaviourSetting
Tool throws a RuntimeExceptionits message goes back to the model as the tool resultspring.ai.tools.throw-exception-on-error=true throws instead
Checked exceptions and Errorsre-thrown—
Calls to one toolat most 40spring.ai.tools.limits.max-calls-per-tool-default
Tool calls in totalat most 150spring.ai.tools.limits.max-total-tool-calls
Limit exceededthrowsspring.ai.tools.limits.on-limit-exceeded (THROW)
returnDirect toolresult returned to the caller; the model never sees ithonoured only if all tools called in that round are returnDirect

Otherwise, the tool result is added to the conversation and the loop goes round again: retrieval, then the model.

The way back

The response travels back up the stack. The memory advisor saves the final answer. If you asked for entity(...), a BeanOutputConverter turns the text into your Java type; format instructions were added to the prompt earlier. The docs call this best effort, so text that doesn’t fit the type fails with an exception, and valid JSON should still be validated before it causes a side effect. Any failure along the way ends as an exception in your service, and your error handling decides the HTTP status. The next request starts from the top, with memory loaded again.

Pause & Prove

The pinned question

Your ChatClient has the memory advisor and the RAG advisor at their default orders, and one tool. The model asks for a tool in two separate rounds, then answers. How many times is the vector store searched, and how many times is memory loaded?

Three searches, one memory load. Two tool rounds plus the final answer make three model calls. The retrieval advisor (order 0) is inside the loop, so it runs once per model call. The memory advisor (HIGHEST_PRECEDENCE + 200) is outside the loop, so it runs once.

The Community poll: when does Spring AI actually contact the model?

  • At call(). The tempting answer. call() only chooses a blocking call over streaming.
  • When you ask for content() or entity(). ✓ Asking for the result (also chatResponse() or responseEntity()) runs the advisor chain and the model call.
  • When the ChatClient is built. Building only configures defaults such as advisors and the system message.
  • At application startup. Startup creates beans and, for MCP, connects the clients; no chat request is sent to the model.

Five things to remember

  1. call() sends nothing; content() or entity() does.
  2. Memory runs once; retrieval runs on every model call.
  3. The model asks for tools; your code runs and authorizes them.
  4. Tool errors go back to the model; loops are capped (40 per tool, 150 in total).
  5. Only the user message and the final answer are saved to memory.

The peg drill

ChatClient: the assistant. Memory: the notes folder. RAG: cards clipped to the brief. Tool: your staff do the lookup. MCP: a standard plug.

Before and after this video

Sources

Spring AI pages show “Spring AI 2.0.1”; all were read on 1 October 2026.

Change notes

  • 1 Oct 2026: first published.
  • Correction (1 Oct 2026): the video says the memory advisor saves the user message and the final answer “on the way out”. In the v2.0.1 source, the user message is saved on the way in (before(), right after the history is loaded) and only the final answer on the way out (after()). What is stored is the same; the timing differs, so a request that fails after the memory advisor has run still leaves the user message in memory.
  • Version note: the MCP client here is the one in Spring AI 2.0.1, which uses the initialize handshake of MCP spec 2025-11-25. The current MCP documentation (protocol 2026-07-28) describes a stateless protocol with a server/discover request instead. The video is correct for 2.0.1; re-check this step when Spring AI adopts the newer revision.

Not affiliated with or endorsed by VMware Broadcom (Spring) or the Model Context Protocol project. Found a mistake? Tell us in the video’s comments and we’ll correct this page.

Found this useful? The deep version lives on YouTube — new breakdowns of how AI dev tools actually work, weekly.

Subscribe on YouTube →

Prefer email? Get the free newsletter: one failure, traced step by step, about once a week.