Skip to main content
Repeat Agent calls resend the same large system prompt every time. Set prompt_caching=True and Swarms marks the stable prefix so the provider reuses it — you pay full price once, then a discount on every call after.

When to use it

  • A large, reusable system prompt (persona, policies, examples) sent on every call.
  • The same agent runs many times or holds a multi-turn conversation.
  • Tool-heavy agents whose tool schemas stay constant across calls.
  • Long context (docs, transcripts) reused turn after turn.

Basic usage

Flip on prompt_caching. On Anthropic, Swarms adds cache_control breakpoints to the stable prefix; on OpenAI, caching is automatic and the flag leaves messages untouched. Caching only kicks in above the provider’s token minimum (see the table in the guide), so the system prompt must be large — here we repeat a string to cross that bar.
  1. prompt_caching=True marks the system prompt (plus the last message) as cacheable.
  2. The first run pays to write the cache.
  3. Every later run reads the cached prefix instead of re-billing it.

Tune it with cache_config

Pass a cache_config dict to control caching. The common knob is ttl — Anthropic supports a 1-hour cache:
cache_config also accepts cache_system_prompt, cache_messages, cache_tools, override, and OpenAI’s prompt_cache_key / prompt_cache_retention. See the full Prompt Caching guide under Agent Development for the complete list.

Verify it worked

agent.usage adds up the provider’s token counts for every call the agent makes. Its cached_tokens key is the part of input_tokens served from the cache. Snapshot it around the second run:
A non-zero cached count on the second run confirms the cached prefix was reused. See Token usage for the full set of fields.

See also

  • Prompt Caching — Full reference: all cache_config keys, provider behavior, and cost details.
  • Context Compression — Shrink long transcripts before they hit the context limit.