Spring AI Words Explained: ChatClient, Advisors, Chat Memory, RAG, Tools, MCP
▶ Watch on YouTube & subscribe to The Stack Underflow
A customer of an online bookstore types: where is my order 42, and can I return the book? To answer, a Spring Boot app has to remember the conversation, look up the returns policy, check the order, and ask a shipping service. The next video in this series traces that one request step by step. This page gives you the words first, so none of them trip you up there.
The one-line version: the model remembers nothing and has never seen your files, so everything Spring AI does (memory, retrieval, tools) comes down to putting the right text into one prompt, and running the tools the model asks for in your own code.
Last verified against the Spring AI 2.0.1 reference documentation and modelcontextprotocol.io: 1 October 2026. The video targets Spring AI 2.0.1 on Spring Boot 4; the Spring AI getting-started page says “Spring AI 2.0.x supports Spring Boot 4.0.x and 4.1.x.” Behaviour changes between versions, so check the docs for yours.
One picture: a consultant with no memory
The video uses one analogy for every word. You hire a brilliant consultant you can phone, but they remember nothing between calls and have never seen your files. So on each call your assistant puts a briefing pack on their desk, with notes from earlier calls and cards from your library, and your staff do any lookups they ask for.
This is a teaching device, not a description of the software. Each word below gets a short peg from that picture, and where the peg stops matching, we say so.
The pieces: Spring AI, the model, tokens
| Word | Memory peg | What it actually is |
|---|---|---|
| Spring AI | the office setup (phone line and staff) | A Spring project for calling AI models. Starters plus Boot auto-configuration create ordinary beans (a ChatModel, a prototype ChatClient.Builder) in your one application context. Not a model, not a second runtime. |
| Model | the consultant with no memory | The AI itself, running at a provider. The docs: “Large language models (LLMs) are stateless, meaning they do not retain information about previous interactions.” |
| Token | billed by the word | The unit a model reads and writes. The docs say one token is roughly 75% of a word in English, and both input and output count toward charges. The peg breaks here: billing is per token, not per word. |
| Context window | the desk | The token limit for one call. It is the size of one request, not memory: history, documents and tool results all have to fit on it. |
Talking to the model
| Word | Memory peg | What it actually is |
|---|---|---|
| Prompt and roles | the briefing pack | A list of messages plus chat options. System guides behaviour, user is the question, assistant is the model’s reply, tool carries the result of a tool the model asked for. |
ChatModel | the phone line | The lower-level, portable interface: call(Prompt) returns a ChatResponse. In Spring AI 2 it sends tool definitions and returns the model’s response; it does not run your tools. |
ChatClient | the assistant | The fluent API you actually use: prompt(), then system(), user(), advisors(), then call(), then content() or entity(). |
| Advisor, advisor chain | the checklist (a stack) | Advisors intercept a ChatClient call to modify the request on the way in and the response on the way out. |
Two details matter later. First, call() alone sends nothing. The ChatClient reference says: “Calling the call() method does not actually trigger the AI model execution.” The model is contacted when you ask for the result. Second, the advisor chain is a stack: the advisor with the lowest order value “will be the first to process the request” and “the last to process the response”. Advisors are Spring AI’s own interfaces set on the ChatClient, not Spring AOP advice. ChatClient also auto-registers a ToolCallingAdvisor unless you disable that.
String answer = chatClient.prompt()
.system("You answer bookstore support questions.")
.user(question)
.advisors(a -> a.param(ChatMemory.CONVERSATION_ID, conversationId))
.call() // nothing sent yet
.content(); // the model is called here
Remembering and looking things up
Both mechanisms in this section work the same way: they add text to the prompt.
| Word | Memory peg | What it actually is |
|---|---|---|
| Chat memory | the notes folder, put back on the desk each call | Recent messages re-sent with each prompt. The default MessageWindowChatMemory keeps the last 20 messages and always keeps system messages. |
| Chat history | the archive | Every message ever exchanged. ChatMemory is not designed to be the archive; the docs point to Spring Data for a complete record. |
ChatMemoryRepository | the filing cabinet | Its “sole responsibility is to store and retrieve messages”. The default keeps them in a ConcurrentHashMap, so they are gone when the app restarts; JDBC and other repositories exist. The 20-message window is a memory rule, not a storage rule. |
| Conversation ID | the case number | The key for one conversation, which you choose (not a user ID). In Spring AI 2.0.1 it is required on every memory call: omitting it throws IllegalArgumentException, and “There is no default conversation ID.” |
| Embedding | a spot on the meaning map | Text turned into an array of floating point numbers designed to capture meaning; close vectors mean similar texts. Not encryption and not a hash. A real embedding has many dimensions, not two. |
EmbeddingModel | the mapmaker | The interface that converts text to vectors, here by calling the provider. |
| Vector store | the library shelved by meaning | Saves vectors with their text and metadata, and does similarity search instead of exact matches. Not your database of record: order 42 lives in your orders database. |
| ETL, chunk | cutting index cards (then shelving them) | A DocumentReader reads a source, a splitter such as TokenTextSplitter cuts it into chunks by token count, and the VectorStore (itself a DocumentWriter) embeds and saves them. ETL runs when you decide, not inside a chat request. |
| RAG | cards clipped to the brief | Retrieval augmented generation: embed the question, find the closest chunks, add them to the user message as context. Nothing is trained. By default, if nothing is found, the model is told not to answer. |
Letting the model act: tools and MCP
| Word | Memory peg | What it actually is |
|---|---|---|
Tool, @Tool | your staff do the lookup | A method the model may ask for. The tools reference: “The model can request a tool call and provide input arguments; the application is responsible for executing the tool and returning the result.” The model never runs your code or gets direct access to your APIs. |
ToolContext | (no peg) | Data for your tool, such as a tenant ID, that is “not sent to the model”. |
| MCP | a standard plug | The Model Context Protocol: a standard protocol that lets an AI application use tools, resources and prompts offered by another program. Not a model and not a Spring library; SDKs exist for several languages. |
| MCP client and server | the socket and the device | The host (your app) creates one MCP client for each server it connects to. Spring AI wraps the server’s tools as tool callbacks, so they look like local tools; a tools/call request runs the tool in the server’s own process. |
Where the “staff” peg breaks: with Spring AI defaults, the auto-registered ToolCallingAdvisor executes requested tools automatically. “Your staff decide” means your code’s design (what the tool checks before acting), not a human approving each call.
Getting the answer out
| Word | Memory peg | What it actually is |
|---|---|---|
| Structured output | fill in the form | With entity(...), a converter such as BeanOutputConverter adds format instructions (a JSON schema) to the prompt and parses the reply into your Java type. The docs call this best effort: “The AI Model is not guaranteed to return the structured output as requested.” Validate before acting on it. |
| Streaming | hear it live | stream() instead of call(); content() then returns a Flux of strings as the model generates them. These pieces are parts of the answer, not RAG chunks. |
Pause & Prove
The pinned question: chat memory and RAG both add text to the prompt. What does each add, and where does it come from?
- Chat memory adds this conversation’s recent messages, loaded from the
ChatMemoryby conversation ID: by default the last 20, with system messages always kept. - RAG adds document chunks, retrieved from the vector store by similarity of meaning to the question.
- Neither trains the model. Both exist only in this one prompt; the next call starts from an empty desk again.
The Community poll: your Spring AI app gives the model a tool. Who actually runs it?
- The model. The tempting answer. The model only sends back a tool name and arguments; it never executes your code.
- The model provider. The provider hosts the model, not your tool. Your method’s code and your APIs never leave your side.
- Your application (or an MCP server). ✓ A local
@Toolmethod runs in your app; an MCP tool runs in the MCP server’s own process, reached through your app’s MCP client. - The vector store. The vector store answers similarity searches for RAG. It doesn’t execute tools.
The peg drill
Say the peg for each, then check: context window: the desk. RAG: cards clipped to the brief. MCP: a standard plug.
Before and after this video
- Next: What Spring AI Does Between chatClient.prompt() and the Answer follows the bookstore question through memory, the tool loop, retrieval and the MCP call as one state machine.
- Further reading: Spring AI ChatClient Explained from our Spring AI for Enterprise Java series, on system vs. user messages and reading token usage from
ChatResponseMetadata. - Earlier in this series: Spring Boot words: bean, ApplicationContext, auto-configuration, the words this page assumes.
Sources
Spring AI pages show “Spring AI 2.0.1”; all were read on 1 October 2026.
- Getting started: 2.0.1 is the current GA; 2.0.x supports Spring Boot 4.0.x and 4.1.x
- AI concepts: models, tokens (~75% of a word, input and output billed), context window, embeddings
- ChatClient API: prototype
ChatClient.Builderbean,call()does not trigger the model,ToolCallingAdvisorauto-registered,stream().content()returns aFlux - Advisors API: intercept and modify; stack order
- Chat memory: stateless models, memory vs history, window of 20, repository role, required conversation ID
- ETL pipeline: reader, transformer, writer;
TokenTextSplitter;VectorStoreasDocumentWriter - Vector databases: similarity search instead of exact match
- Retrieval augmented generation: retrieved context added to the user text; empty context by default tells the model not to answer
- Tool calling: model requests, application executes;
ToolContextnot sent to the model;ChatModeldoes not execute tools in 2.x - MCP client Boot starter: MCP tools exposed as tool callbacks
- Structured output converters: format instructions, JSON schema, best effort
- MCP architecture overview: one client per server; tools, resources and prompts;
tools/call; SDKs in several languages
Change notes
- 1 Oct 2026: first published. One line from the video is not repeated here: that Spring AI versions before 2.0 fell back to a shared default conversation ID. The current reference only says “There is no default conversation ID”, and we could not re-check the older behaviour on a vendor page today.
- 1 Oct 2026, correction: the video says the memory advisor loads the memory before the call and “saves the new exchange after”. In Spring AI 2.0.1 it saves your new message before the model call and only the answer after (
MessageChatMemoryAdvisor.before()andafter()in the v2.0.1 source). So a request that fails still leaves your message in memory.
Not affiliated with or endorsed by VMware Broadcom (Spring) or the Model Context Protocol project. Found a mistake? Tell us in the video’s comments and we’ll correct this page.
Found this useful? The deep version lives on YouTube — new breakdowns of how AI dev tools actually work, weekly.
Subscribe on YouTube →Prefer email? Get the free newsletter: one failure, traced step by step, about once a week.