Skip to main content
Swarms agents can keep disk-backed persistent memory through the Conversation class. When persistent_memory=True, each agent writes its active interaction log to a MEMORY.md file and reloads that file when another process starts an agent with the same agent_name. The flag is False by default, so agents are ephemeral and keep no on-disk state unless you opt in. Use this page to understand:
  • How MEMORY.md is created, loaded, and updated
  • How context compression keeps memory within the model context window
  • How archived transcripts preserve raw chat history
  • How to inspect, compact, export, or disable memory in code
  • How to give an agent external knowledge with a retrieval tool
Persistent memory is keyed by agent_name. Reusing the same agent_name (with persistent_memory=True) resumes the same memory across process restarts. Changing the name starts a separate memory folder. Leaving persistent_memory at its default of False keeps the agent fully ephemeral.

The persistent_memory flag

persistent_memory is the top-level switch that controls whether the agent reads from and writes to MEMORY.md.
bool
default:"False"
Enables disk-backed persistent memory. When True, the agent creates MEMORY.md on first run, preloads it on subsequent runs, and writes new turns through to disk. When False (the default), the agent runs fully in-process — no MEMORY.md, no archive/, fresh state every run.

Memory stack

An agent can use several memory layers at the same time: MEMORY.md records the agent’s own interaction history. To let an agent look things up in external documents or a database, give it a retrieval tool instead.

Disk layout

Agent memory lives under the workspace directory:
str
default:"swarm-worker-01"
Stable name used to identify the agent’s memory folder.
Set this explicitly when using persistent_memory. An omitted name defaults to the same literal string, "swarm-worker-01", for every agent — so multiple default-named agents don’t just lose history on restart, they all read and write the same MEMORY.md folder at the same time, corrupting each other’s memory.

Key design points

  • The folder is keyed by agent_name, not by id.
  • MEMORY.md is append-updated during normal operation.
  • Every conversation.add(role, content) writes to in-memory history and to disk.
  • Compression archives the current MEMORY.md before replacing it with a compact summary.
  • The agent’s static system_prompt, rules, and constructor configuration are not repeatedly appended to MEMORY.md.

Lifecycle

1. File creation

On first construction of an agent with a new agent_name, Swarms creates:
The file starts with a small header and an interaction log section:
If the file already exists, Swarms leaves it in place.

2. Preload on construction

During Conversation.__init__, Swarms reads the existing MEMORY.md and injects it into conversation_history as a single System message. The resulting prompt order is:
The preload is added directly to memory, so it is not written back to disk again. When return_history_as_string() builds the prompt, the model sees the system prompt, rules, persistent memory, and current task in order.

3. Write-through on new messages

Every conversation.add(role, content) call:
  1. Appends the message to conversation_history
  2. Appends a timestamped block to MEMORY.md
The on-disk format looks like this:
Disk writes are serialized with a per-conversation lock. Construction-time messages such as system prompts and rules are suppressed from disk so static identity does not get duplicated on every restart.

Context compression

Without compression, a long-running agent could eventually exceed the model’s context window. Swarms can attach a ContextCompressor that summarizes the current transcript and compacts the active memory.
bool
default:"True"
Enables automatic compression when memory approaches the configured context limit.

When compression runs

Compression can run when all of these are true:
  • context_compression=True
  • context_length is set to a non-zero value
  • The token usage of short_memory.return_history_as_string() is greater than or equal to threshold * context_length
  • The agent is at the top of a loop iteration
The default threshold is 0.9, so compression starts when the active prompt reaches about 90% of the context window.
context_length accepts None as its constructor default, but the agent never runs with an actual None value: during __init__, an unset context_length is resolved to a model-derived value (or 16000 as a fallback), so compression always has a real denominator to measure against. Pass context_length explicitly only to override that resolved default — for example, to match a larger context window.
Compression works for both max_loops="auto" and integer max_loops runs. The context_compression flag is the gate.

What compression does

When compression fires:
  1. The current transcript is summarized with an LLM call.
  2. Conversation.compact(summary=...) is called.
  3. The current MEMORY.md is copied to archive/history_<timestamp>.md.
  4. The active MEMORY.md is deleted and recreated with a fresh header.
  5. conversation_history is rebuilt with the system prompt, rules, and custom rules.
  6. The summary is appended as one System message to both memory and MEMORY.md.
After compaction, active memory is small again:
On the next process restart, Swarms loads the compact summary from MEMORY.md instead of the raw pre-compaction transcript. The archive keeps the full transcript available without filling the active context window.

Configure compression

Compression is enabled by default:
Disable compression when you want the active MEMORY.md to remain un-compacted:
When compression is enabled, the agent attaches a ContextCompressor(threshold=0.9). You can replace it after construction to tune the threshold, summarizer model, temperature, or summary length:

Access memory in code

The Conversation object is available as agent.short_memory.

Manual compaction

You can compact memory yourself at any time:
Manual compaction follows the same archive, wipe, and re-seed flow as automatic compression.

Export and load conversations

MEMORY.md is the active persistent memory file. You can also export or load conversation history in other formats:

Search memory

Use built-in search helpers for quick inspection:

Disable disk-backed memory

persistent_memory=False is the default, and it keeps an agent fully in-process. Nothing is preloaded, nothing is written to MEMORY.md, and no archive/ directory is created:
If you have already constructed a persistent agent and want to stop further disk writes for the rest of the run, you can also clear memory_md_path:
This stops future writes but does not retroactively delete MEMORY.md. Neither approach disables conversation_history — that always tracks the current run in memory.

External knowledge

MEMORY.md holds the agent’s own interaction history. When the agent needs to look things up in documents, a database, or any other external store, give it a retrieval tool. The agent calls the tool during its loop and the results enter the conversation like any other tool output.
Swap the dictionary lookup for a real vector store, SQL query, or HTTP call and the shape of the code stays the same. See Agent Tools for the full tool interface.
Union[Callable, Any]
default:"None"
An arbitrary object attached to the agent and stored on agent.long_term_memory.
agent.run() never queries this object. The only method Swarms calls on it is .save(path), and only when autosave writes agent state to disk. There is no automatic retrieval in the run loop and no built-in vector database ships with Swarms. Use a retrieval tool, as shown above, for anything the agent must actually read during a run.

Best practices

  • Use stable, descriptive agent_name values for agents that should remember previous work.
  • Keep context_compression=True and set context_length for autonomous or long-running agents.
  • Tune ContextCompressor.threshold lower for agents with large tool outputs or long responses.
  • Compact manually after major milestones to preserve the important state and reduce prompt size.
  • Use a retrieval tool for external knowledge. Do not rely on MEMORY.md as a document database.
  • Leave persistent_memory at its default False for privacy-sensitive or one-off agents that should not write a transcript.

Why it works this way

Why key memory by agent_name?

id values can change between process starts. agent_name is user-controlled and stable, so it gives the agent a durable identity.

Why preload memory as one System message?

The model needs to understand that the content is prior memory, not a current user request. A single system-level memory preamble is compact and less ambiguous than replaying old turns as active messages.

Why wipe MEMORY.md during compaction?

If compaction only appended a summary, the next run would load both the summary and the raw transcript it summarizes. Wiping the active file keeps the working context small, while archive/ preserves the raw log.

Next steps

Agent Configuration

Configure core agent parameters such as agent_name, max_loops, and context limits.

Conversation API

Explore the underlying Conversation class and its export, load, and search helpers.