Skip to main content
Swarms gives you two token numbers for an agent:
  • agent.usage is what the provider billed. It sums the token counts the provider reports on every call the agent has made.
  • agent.input_tokens is an estimate of the next request. It counts what the agent is about to send, before you run it.
SwarmRouter, GraphWorkflow, HeavySwarm and TreeOfThoughts expose a usage property with the same keys, summed over the agents they run.

agent.usage

Fields

agent.usage returns a new dict on every access, so changing it does not change the agent’s totals.
int
Prompt tokens, from the provider’s prompt_tokens.
int
Completion tokens, from the provider’s completion_tokens.
int
The part of input_tokens the provider served from its prompt cache, from prompt_tokens_details.cached_tokens. It is already included in input_tokens, not added to it. See Prompt Caching.
int
The part of output_tokens the model spent on hidden reasoning, from completion_tokens_details.reasoning_tokens. It is already included in output_tokens.
int
The provider’s total_tokens, or input_tokens + output_tokens when the provider leaves it out.
LiteLLM reports every provider in OpenAI’s usage shape, so the keys are the same for OpenAI, Anthropic and the rest. A breakdown the provider does not report counts as 0, which means unknown, not zero.

What is counted

Every completion the agent’s own model client makes adds to the total:
  • Every loop of run(), including turns that only call tools.
  • The tool-summary call after a tool runs (tool_call_summary=True, the default) and the summary after MCP tool calls.
  • The plan, execute and summary calls of max_loops="auto", and the planning call of plan_enabled=True.
  • Calls made after switching to a fallback model. The total lives on the agent, so rebuilding the model client does not reset it.
  • Streamed calls. Swarms asks the provider to end the stream with a usage chunk and records it when the stream has been consumed. run() consumes its own streams; with run_stream() and arun_stream() the total updates once you have read every token.
A response without a usage block, such as a failed request, adds nothing. Calls made by other agents are not counted, even when this agent started them. Sub-agents created in autonomous mode and agents reached through handoffs keep their own usage. Sum them yourself:
If you pass your own LiteLLM instance as llm, the agent sets its usage_hook so its calls land in agent.usage too, unless you set a hook yourself.

When it resets

Never. agent.usage is a lifetime total for the agent instance: it keeps growing across runs and there is no reset method. save() and load() do not store it, so a loaded agent starts from zero. To measure one run, snapshot before and after:
Read agent.usage, not agent.llm.usage. The model client keeps its own counter, but the agent replaces the client when it switches to a fallback model or loads tools on demand, and the new client starts from zero.

Estimate cost

Multiply by your provider’s rates. Cached input is usually billed at a lower rate, so price it separately:

Tool loops

Each loop that calls a tool makes two provider calls, the tool call and the summary of its result. Both are counted:

agent.input_tokens

agent.input_tokens counts the tokens the agent’s next request would carry, using its model’s tokenizer. Use it to see how full the context window is before a run:
It covers the system prompt, the whole conversation in agent.short_memory, and the tool schemas in agent.tools_list_dictionary. It does not cover text added only at request time, such as the Agent Skills section or MCP tool schemas.
input_tokens is an estimate. The conversation is counted as rendered text, with role labels and timestamps, so it lands a little above what the provider bills for the same content. It makes no provider call and is computed fresh on every access. For billed numbers, use usage.

Usage across a swarm

Every structure below returns a new dict with the same five keys, built by adding up Agent.usage. Because agent totals are lifetime totals, the structure totals grow across runs too, and an agent shared with another workflow contributes what it spent there as well. Snapshot before and after a run to isolate it.

SwarmRouter.usage

router.usage sums Agent.usage over the agents you passed in, plus every Agent that a swarm the router has built holds as an attribute, directly or in a list or tuple. That covers agents the swarm creates for itself, such as a HierarchicalSwarm director. Each agent is counted once.
router.usage only finds agents stored as attributes of the built swarm, alone or in a list or tuple. It misses agents a swarm keeps anywhere else. With swarm_type="HeavySwarm", the workers live in a dict and the question generator is not an Agent, so router.usage reports only the agents you passed in. Read router.swarm.usage after the run instead. router.swarm is the swarm that ran last.

GraphWorkflow.usage

workflow.usage sums Agent.usage over every agent node, including the agents inside nested subgraph workflows at any depth. An agent that backs several nodes, or appears in both the graph and a subgraph, is counted once. Nodes without an agent are skipped.

HeavySwarm.usage

swarm.usage adds the question-generation calls, which go through a bare LiteLLM client rather than an Agent, to Agent.usage for every agent in swarm.agents: the workers plus the synthesis agent or captain, depending on variant. It grows across run() calls and across max_loops.

TreeOfThoughts.usage

TreeOfThoughts runs each model call on a single-use agent and folds its usage into the search. tot.usage is the total over every search the instance has run. tot.last_result.usage holds the latest search alone, next to last_result.llm_calls and last_result.nodes_expanded. See Tree of Thoughts.

Other structures

SequentialWorkflow, ConcurrentWorkflow, AgentRearrange and the other structures do not expose usage. Sum it over the agents you gave them:
Or run them through a SwarmRouter and read router.usage.

Summary

Reference

  • Agent.usage, Agent.input_tokens: swarms/structs/agent.py
  • Usage normalization and stream accounting: empty_usage, usage_from_response and LiteLLM._track_streaming_usage in swarms/utils/litellm_wrapper.py
  • SwarmRouter.usage: swarms/structs/swarm_router.py
  • GraphWorkflow.usage: swarms/structs/graph_workflow.py
  • HeavySwarm.usage: swarms/structs/heavy_swarm.py
  • TreeOfThoughts.usage: swarms/agents/tree_of_thoughts.py
  • Runnable examples: examples/single_agent/utils/agent_usage.py and examples/multi_agent/swarm_router/swarm_router_usage.py