agent.usageis what the provider billed. It sums the token counts the provider reports on every call the agent has made.agent.input_tokensis an estimate of the next request. It counts what the agent is about to send, before you run it.
SwarmRouter, GraphWorkflow, HeavySwarm and TreeOfThoughts expose a usage property with the same keys, summed over the agents they run.
agent.usage
Fields
agent.usage returns a new dict on every access, so changing it does not change the agent’s totals.
int
Prompt tokens, from the provider’s
prompt_tokens.int
Completion tokens, from the provider’s
completion_tokens.int
The part of
input_tokens the provider served from its prompt cache, from prompt_tokens_details.cached_tokens. It is already included in input_tokens, not added to it. See Prompt Caching.int
The part of
output_tokens the model spent on hidden reasoning, from completion_tokens_details.reasoning_tokens. It is already included in output_tokens.int
The provider’s
total_tokens, or input_tokens + output_tokens when the provider leaves it out.LiteLLM reports every provider in OpenAI’s usage shape, so the keys are the same for OpenAI, Anthropic and the rest. A breakdown the provider does not report counts as
0, which means unknown, not zero.What is counted
Every completion the agent’s own model client makes adds to the total:- Every loop of
run(), including turns that only call tools. - The tool-summary call after a tool runs (
tool_call_summary=True, the default) and the summary after MCP tool calls. - The plan, execute and summary calls of
max_loops="auto", and the planning call ofplan_enabled=True. - Calls made after switching to a fallback model. The total lives on the agent, so rebuilding the model client does not reset it.
- Streamed calls. Swarms asks the provider to end the stream with a usage chunk and records it when the stream has been consumed.
run()consumes its own streams; withrun_stream()andarun_stream()the total updates once you have read every token.
handoffs keep their own usage. Sum them yourself:
LiteLLM instance as llm, the agent sets its usage_hook so its calls land in agent.usage too, unless you set a hook yourself.
When it resets
Never.agent.usage is a lifetime total for the agent instance: it keeps growing across runs and there is no reset method. save() and load() do not store it, so a loaded agent starts from zero. To measure one run, snapshot before and after:
Estimate cost
Multiply by your provider’s rates. Cached input is usually billed at a lower rate, so price it separately:Tool loops
Each loop that calls a tool makes two provider calls, the tool call and the summary of its result. Both are counted:agent.input_tokens
agent.input_tokens counts the tokens the agent’s next request would carry, using its model’s tokenizer. Use it to see how full the context window is before a run:
agent.short_memory, and the tool schemas in agent.tools_list_dictionary. It does not cover text added only at request time, such as the Agent Skills section or MCP tool schemas.
input_tokens is an estimate. The conversation is counted as rendered text, with role labels and timestamps, so it lands a little above what the provider bills for the same content. It makes no provider call and is computed fresh on every access. For billed numbers, use usage.Usage across a swarm
Every structure below returns a new dict with the same five keys, built by adding upAgent.usage. Because agent totals are lifetime totals, the structure totals grow across runs too, and an agent shared with another workflow contributes what it spent there as well. Snapshot before and after a run to isolate it.
SwarmRouter.usage
router.usage sums Agent.usage over the agents you passed in, plus every Agent that a swarm the router has built holds as an attribute, directly or in a list or tuple. That covers agents the swarm creates for itself, such as a HierarchicalSwarm director. Each agent is counted once.
GraphWorkflow.usage
workflow.usage sums Agent.usage over every agent node, including the agents inside nested subgraph workflows at any depth. An agent that backs several nodes, or appears in both the graph and a subgraph, is counted once. Nodes without an agent are skipped.
HeavySwarm.usage
swarm.usage adds the question-generation calls, which go through a bare LiteLLM client rather than an Agent, to Agent.usage for every agent in swarm.agents: the workers plus the synthesis agent or captain, depending on variant. It grows across run() calls and across max_loops.
TreeOfThoughts.usage
TreeOfThoughts runs each model call on a single-use agent and folds its usage into the search. tot.usage is the total over every search the instance has run. tot.last_result.usage holds the latest search alone, next to last_result.llm_calls and last_result.nodes_expanded. See Tree of Thoughts.
Other structures
SequentialWorkflow, ConcurrentWorkflow, AgentRearrange and the other structures do not expose usage. Sum it over the agents you gave them:
SwarmRouter and read router.usage.
Summary
Reference
Agent.usage,Agent.input_tokens:swarms/structs/agent.py- Usage normalization and stream accounting:
empty_usage,usage_from_responseandLiteLLM._track_streaming_usageinswarms/utils/litellm_wrapper.py SwarmRouter.usage:swarms/structs/swarm_router.pyGraphWorkflow.usage:swarms/structs/graph_workflow.pyHeavySwarm.usage:swarms/structs/heavy_swarm.pyTreeOfThoughts.usage:swarms/agents/tree_of_thoughts.py- Runnable examples:
examples/single_agent/utils/agent_usage.pyandexamples/multi_agent/swarm_router/swarm_router_usage.py