Spring AI Words Explained: ChatClient, Advisors, Chat Memory, RAG, Tools, MCP

October 1, 2026 · How It Actually Works: Real Systems as State Machines (part 8)

▶ Watch on YouTube & subscribe to The Stack Underflow

A customer of an online bookstore types: where is my order 42, and can I return the book? To answer, a Spring Boot app has to remember the conversation, look up the returns policy, check the order, and ask a shipping service. The next video in this series traces that one request step by step. This page gives you the words first, so none of them trip you up there.

The one-line version: the model remembers nothing and has never seen your files, so everything Spring AI does (memory, retrieval, tools) comes down to putting the right text into one prompt, and running the tools the model asks for in your own code.

Last verified against the Spring AI 2.0.1 reference documentation and modelcontextprotocol.io: 1 October 2026. The video targets Spring AI 2.0.1 on Spring Boot 4; the Spring AI getting-started page says “Spring AI 2.0.x supports Spring Boot 4.0.x and 4.1.x.” Behaviour changes between versions, so check the docs for yours.

One picture: a consultant with no memory

The video uses one analogy for every word. You hire a brilliant consultant you can phone, but they remember nothing between calls and have never seen your files. So on each call your assistant puts a briefing pack on their desk, with notes from earlier calls and cards from your library, and your staff do any lookups they ask for.

This is a teaching device, not a description of the software. Each word below gets a short peg from that picture, and where the peg stops matching, we say so.

The pieces: Spring AI, the model, tokens

WordMemory pegWhat it actually is
Spring AIthe office setup (phone line and staff)A Spring project for calling AI models. Starters plus Boot auto-configuration create ordinary beans (a ChatModel, a prototype ChatClient.Builder) in your one application context. Not a model, not a second runtime.
Modelthe consultant with no memoryThe AI itself, running at a provider. The docs: “Large language models (LLMs) are stateless, meaning they do not retain information about previous interactions.”
Tokenbilled by the wordThe unit a model reads and writes. The docs say one token is roughly 75% of a word in English, and both input and output count toward charges. The peg breaks here: billing is per token, not per word.
Context windowthe deskThe token limit for one call. It is the size of one request, not memory: history, documents and tool results all have to fit on it.

Talking to the model

WordMemory pegWhat it actually is
Prompt and rolesthe briefing packA list of messages plus chat options. System guides behaviour, user is the question, assistant is the model’s reply, tool carries the result of a tool the model asked for.
ChatModelthe phone lineThe lower-level, portable interface: call(Prompt) returns a ChatResponse. In Spring AI 2 it sends tool definitions and returns the model’s response; it does not run your tools.
ChatClientthe assistantThe fluent API you actually use: prompt(), then system(), user(), advisors(), then call(), then content() or entity().
Advisor, advisor chainthe checklist (a stack)Advisors intercept a ChatClient call to modify the request on the way in and the response on the way out.

Two details matter later. First, call() alone sends nothing. The ChatClient reference says: “Calling the call() method does not actually trigger the AI model execution.” The model is contacted when you ask for the result. Second, the advisor chain is a stack: the advisor with the lowest order value “will be the first to process the request” and “the last to process the response”. Advisors are Spring AI’s own interfaces set on the ChatClient, not Spring AOP advice. ChatClient also auto-registers a ToolCallingAdvisor unless you disable that.

String answer = chatClient.prompt()
    .system("You answer bookstore support questions.")
    .user(question)
    .advisors(a -> a.param(ChatMemory.CONVERSATION_ID, conversationId))
    .call()      // nothing sent yet
    .content();  // the model is called here

Remembering and looking things up

Both mechanisms in this section work the same way: they add text to the prompt.

WordMemory pegWhat it actually is
Chat memorythe notes folder, put back on the desk each callRecent messages re-sent with each prompt. The default MessageWindowChatMemory keeps the last 20 messages and always keeps system messages.
Chat historythe archiveEvery message ever exchanged. ChatMemory is not designed to be the archive; the docs point to Spring Data for a complete record.
ChatMemoryRepositorythe filing cabinetIts “sole responsibility is to store and retrieve messages”. The default keeps them in a ConcurrentHashMap, so they are gone when the app restarts; JDBC and other repositories exist. The 20-message window is a memory rule, not a storage rule.
Conversation IDthe case numberThe key for one conversation, which you choose (not a user ID). In Spring AI 2.0.1 it is required on every memory call: omitting it throws IllegalArgumentException, and “There is no default conversation ID.”
Embeddinga spot on the meaning mapText turned into an array of floating point numbers designed to capture meaning; close vectors mean similar texts. Not encryption and not a hash. A real embedding has many dimensions, not two.
EmbeddingModelthe mapmakerThe interface that converts text to vectors, here by calling the provider.
Vector storethe library shelved by meaningSaves vectors with their text and metadata, and does similarity search instead of exact matches. Not your database of record: order 42 lives in your orders database.
ETL, chunkcutting index cards (then shelving them)A DocumentReader reads a source, a splitter such as TokenTextSplitter cuts it into chunks by token count, and the VectorStore (itself a DocumentWriter) embeds and saves them. ETL runs when you decide, not inside a chat request.
RAGcards clipped to the briefRetrieval augmented generation: embed the question, find the closest chunks, add them to the user message as context. Nothing is trained. By default, if nothing is found, the model is told not to answer.

Letting the model act: tools and MCP

WordMemory pegWhat it actually is
Tool, @Toolyour staff do the lookupA method the model may ask for. The tools reference: “The model can request a tool call and provide input arguments; the application is responsible for executing the tool and returning the result.” The model never runs your code or gets direct access to your APIs.
ToolContext(no peg)Data for your tool, such as a tenant ID, that is “not sent to the model”.
MCPa standard plugThe Model Context Protocol: a standard protocol that lets an AI application use tools, resources and prompts offered by another program. Not a model and not a Spring library; SDKs exist for several languages.
MCP client and serverthe socket and the deviceThe host (your app) creates one MCP client for each server it connects to. Spring AI wraps the server’s tools as tool callbacks, so they look like local tools; a tools/call request runs the tool in the server’s own process.

Where the “staff” peg breaks: with Spring AI defaults, the auto-registered ToolCallingAdvisor executes requested tools automatically. “Your staff decide” means your code’s design (what the tool checks before acting), not a human approving each call.

Getting the answer out

WordMemory pegWhat it actually is
Structured outputfill in the formWith entity(...), a converter such as BeanOutputConverter adds format instructions (a JSON schema) to the prompt and parses the reply into your Java type. The docs call this best effort: “The AI Model is not guaranteed to return the structured output as requested.” Validate before acting on it.
Streaminghear it livestream() instead of call(); content() then returns a Flux of strings as the model generates them. These pieces are parts of the answer, not RAG chunks.

Pause & Prove

The pinned question: chat memory and RAG both add text to the prompt. What does each add, and where does it come from?

  • Chat memory adds this conversation’s recent messages, loaded from the ChatMemory by conversation ID: by default the last 20, with system messages always kept.
  • RAG adds document chunks, retrieved from the vector store by similarity of meaning to the question.
  • Neither trains the model. Both exist only in this one prompt; the next call starts from an empty desk again.

The Community poll: your Spring AI app gives the model a tool. Who actually runs it?

  • The model. The tempting answer. The model only sends back a tool name and arguments; it never executes your code.
  • The model provider. The provider hosts the model, not your tool. Your method’s code and your APIs never leave your side.
  • Your application (or an MCP server). ✓ A local @Tool method runs in your app; an MCP tool runs in the MCP server’s own process, reached through your app’s MCP client.
  • The vector store. The vector store answers similarity searches for RAG. It doesn’t execute tools.

The peg drill

Say the peg for each, then check: context window: the desk. RAG: cards clipped to the brief. MCP: a standard plug.

Before and after this video

Sources

Spring AI pages show “Spring AI 2.0.1”; all were read on 1 October 2026.

  • Getting started: 2.0.1 is the current GA; 2.0.x supports Spring Boot 4.0.x and 4.1.x
  • AI concepts: models, tokens (~75% of a word, input and output billed), context window, embeddings
  • ChatClient API: prototype ChatClient.Builder bean, call() does not trigger the model, ToolCallingAdvisor auto-registered, stream().content() returns a Flux
  • Advisors API: intercept and modify; stack order
  • Chat memory: stateless models, memory vs history, window of 20, repository role, required conversation ID
  • ETL pipeline: reader, transformer, writer; TokenTextSplitter; VectorStore as DocumentWriter
  • Vector databases: similarity search instead of exact match
  • Retrieval augmented generation: retrieved context added to the user text; empty context by default tells the model not to answer
  • Tool calling: model requests, application executes; ToolContext not sent to the model; ChatModel does not execute tools in 2.x
  • MCP client Boot starter: MCP tools exposed as tool callbacks
  • Structured output converters: format instructions, JSON schema, best effort
  • MCP architecture overview: one client per server; tools, resources and prompts; tools/call; SDKs in several languages

Change notes

  • 1 Oct 2026: first published. One line from the video is not repeated here: that Spring AI versions before 2.0 fell back to a shared default conversation ID. The current reference only says “There is no default conversation ID”, and we could not re-check the older behaviour on a vendor page today.
  • 1 Oct 2026, correction: the video says the memory advisor loads the memory before the call and “saves the new exchange after”. In Spring AI 2.0.1 it saves your new message before the model call and only the answer after (MessageChatMemoryAdvisor.before() and after() in the v2.0.1 source). So a request that fails still leaves your message in memory.

Not affiliated with or endorsed by VMware Broadcom (Spring) or the Model Context Protocol project. Found a mistake? Tell us in the video’s comments and we’ll correct this page.

Found this useful? The deep version lives on YouTube — new breakdowns of how AI dev tools actually work, weekly.

Subscribe on YouTube →

Prefer email? Get the free newsletter: one failure, traced step by step, about once a week.