Exploring the breakthrough innovations shaping our world. From AI infrastructure and robotics to biotech, quantum computing, and spatial tech.
When an AI assistant appears to remember something from an earlier conversation, several different mechanisms may be responsible. Context, memory, and state are closely related, but they solve different problems.
Context is the information available to the model during a particular interaction. Memory refers to information deliberately retained for future use. State is the working information an application needs to keep track of an ongoing conversation, task, or workflow.
The distinction matters because an AI system does not simply “remember everything.” Developers have to decide what information should be available now, what should be retained for later, and what needs to be tracked while a task is running.
The simplest way to understand context is to think of it as the information supplied to the model when it generates a response.
That can include the user's latest message, earlier conversation messages, system instructions, retrieved documents, tool results, and other information assembled by the application.
Suppose a user says, “I prefer short answers,” and then asks another question. If the earlier preference is included in the model's current context, it can influence the response.
But context is not automatically permanent memory. If the application does not provide information from an earlier interaction, the model cannot use that information merely because it appeared in a previous exchange.
Context also has practical limits. Models have finite context windows, and very long histories can increase processing costs and response time. Large amounts of old or irrelevant information can also make it harder for the model to focus on what matters to the current request. LangChain's documentation discusses these challenges as part of its guidance on managing conversation memory.
This creates an important distinction:
Information can be available to a model now without being something the system has chosen to remember permanently.

Memory is the mechanism used to retain selected information beyond the immediate interaction.
Imagine a user tells an assistant that they prefer temperatures in Fahrenheit. If the application saves that preference and retrieves it during a later conversation, the assistant can use it without asking the user again.
AI applications commonly distinguish between short-term and long-term memory. Short-term memory generally covers information associated with a particular conversation or thread, while long-term memory can persist across separate sessions. LangChain's documentation describes long-term memory as information that can be recalled across conversations and threads, while short-term memory is associated with an ongoing thread.
The storage mechanism can vary. An application may use a conventional database, a document store, or another retrieval system. The important point is that persistent memory is typically an application-level storage and retrieval mechanism, rather than the model spontaneously rewriting its own training.
That distinction also makes memory easier to update. A stored preference can be changed or deleted without retraining the underlying language model.
State is broader than memory because it describes the information an application needs to continue operating correctly.
Consider an AI agent working on a research task. It may need to track which sources have already been checked, what information has been retrieved, which step it is currently performing, whether a tool call failed, and whether human approval is still required.
Those details constitute the application's state.
State does not necessarily represent something the system wants to remember forever. The fact that an agent is currently waiting for approval may be useful for the duration of a task but irrelevant after the task has been completed.
LangChain and LangGraph use state to hold short-term information such as conversation history and other custom fields. State can also be persisted through checkpoints so an ongoing thread can be resumed later.

The three concepts become easier to understand when you follow a piece of information through an AI application.
Suppose a user says, “I always want temperatures displayed in Fahrenheit.”
At first, that preference may simply be part of the current conversation context. The application can then decide that it is useful enough to save as long-term memory. During a later conversation, the application retrieves the preference and places it back into the model's current context.
Meanwhile, the application's state may contain information about the current conversation, retrieved documents, intermediate results, pending tool calls, or approval requirements.
In practical terms, context answers “What does the model have available right now?” Memory answers “What information has the application chosen to retain?” State answers “What does the application need to keep track of to continue the current task?”
The boundaries are not absolute. In some frameworks, short-term memory is implemented as part of agent state, while that state can be persisted and later used to reconstruct context. LangGraph's documentation explicitly describes short-term memory as part of an agent's state.
Understanding that relationship is more useful than treating the three terms as completely separate technologies.
At first glance, the easiest way to preserve a conversation seems to be sending the entire history to the model every time. That approach becomes less attractive as the conversation grows.
A large history consumes tokens, can increase latency and cost, and may contain old information that is no longer relevant. Even when a model supports a large context window, sending everything does not necessarily produce the best result.
Developers therefore use several strategies to control what reaches the model.
A sliding window keeps the most recent portion of a conversation and removes older messages from the active context as new messages arrive.
This is relatively simple and can work well when recent exchanges are much more important than distant ones. The drawback is that information that falls outside the window may no longer be available unless the application has stored it somewhere else.
Message trimming is one documented approach to controlling how much conversation history is passed to an LLM.
A summary buffer or similar summarization approach replaces older conversation turns with a condensed description of the important information.
Instead of repeatedly sending dozens of earlier messages, the application can preserve their key points in a shorter form. This can reduce the amount of context while keeping some information from earlier parts of the conversation.
The tradeoff is that summarization can lose details. A fact that seemed unimportant when the summary was created may become relevant later.
For information stored outside the conversation, retrieval-augmented generation (RAG) offers another approach.
Rather than placing an entire collection of documents into the model's context, a retrieval system searches an external knowledge source and provides relevant passages when they are needed. Vector retrieval is one common technique for finding information based on semantic similarity.
LangChain's retrieval documentation describes this approach as a way to provide LLMs with relevant external information without requiring the entire knowledge base to fit inside every model request.
This can be particularly useful for large collections of company documents, technical documentation, product information, or other material that would be impractical to include in every conversation.
RAG and memory, however, are not interchangeable. A user's saved preference might belong in long-term memory, while a relevant section of a large document collection can simply be retrieved when the current question requires it.
Persistence does not make information automatically trustworthy.
A user can change a preference. A company's policy can be replaced. An earlier inference can turn out to be incorrect. If outdated information remains in memory, an assistant may continue using it after it is no longer appropriate.
A useful memory system therefore needs more than storage. It needs ways to update, replace, or delete information.
This becomes particularly important when persistent information relates to people or organizations. Developers need to consider what is stored, how it is retrieved, who can access it, and how long it should remain available.
The ability to retain information does not automatically mean that the information should be retained.
State becomes especially valuable when an AI system performs work over several steps.
An agent researching a question might retrieve several documents, identify a missing piece of information, call another tool, and then wait for a human decision. If those intermediate details are recorded in state, the system can continue from where it stopped rather than reconstructing the entire process.
Persistent checkpoints can also help applications recover after interruptions. LangGraph supports persistence of agent state and checkpoints that can be used to resume ongoing threads.
State can also make an AI system easier to debug. Developers can inspect what the application believed had already happened, which results had been collected, and what step was pending when a problem occurred.
That is especially valuable for agents, where the final answer may not reveal which earlier decision caused the system to go off track.
One common misconception is that an AI assistant has been retrained whenever it appears to remember something.
Usually, that is not what happened.
Training changes a model's parameters based on its training data. An application-level memory system generally works by storing information separately and retrieving it when appropriate.
That separation allows a user's preference to be changed without changing the model itself. It also allows an enterprise assistant to work with current company information without retraining the underlying model every time a document or policy changes.
The tradeoff is that the application now has to manage storage, retrieval, permissions, data quality, and consistency.
For developers, three questions provide a useful starting point.
Context: What does the model need to see right now?
Memory: What information is worth retaining for later?
State: What does the application need to track so the current task can continue correctly?
Keeping those questions separate can prevent unnecessary complexity. A temporary tool result does not necessarily belong in long-term memory. A user's permanent preference does not need to be manually repeated in every conversation. An agent's current approval status may belong in task state without becoming part of the user's long-term profile.
The goal is not to make an AI system remember everything. The better goal is to make sure the right information is available at the right time.
That is the practical relationship among context, memory, and state: context supports the model's immediate decision, memory provides persistence, and state keeps the larger application on track.