Skip to main content

Overview

A decision model answers typed questions about a piece of content, which it calls the state. You send the state and a set of questions in one request. Each answer comes back as structured data with probabilities, not as generated text. DecisionModel is the client for these models. It supports three question types:
TypeSafe’s jev-latest is the default model. Set model_name to "clef" or "clef-flash" to use Cloudflare’s Clef models on Workers AI instead.

When to use it

Use a decision model where you would otherwise call an LLM only to make a discrete judgement and then parse its text:
  • Routing: pick which agent should handle a task, and escalate when confidence is low.
  • Guardrails: score every message for hazards before it is posted.
  • Early stopping: ask whether a debate or a swarm has stopped making progress.
  • Judging: rate many candidate answers against a rubric in one request.
  • Screening: run the same questions over hundreds of items concurrently.
Use an Agent when you need generated text, tool calls or multi-step reasoning.
DecisionModel is a standalone HTTP client built on httpx. It is not an Agent, it does not go through LiteLLM, and no swarm structure takes it as a parameter. You call it from your own code, before, between or after agent runs. See Using it with agents and swarms.

Import

Both names are also exported from swarms.structs. The module is swarms.structs.decision_model.

Providers

The model name picks the provider. A name that starts with clef goes to Cloudflare. Every other name goes to TypeSafe, including names that start with jev and gateway ids such as ~typesafe/jev-latest. Every request is a POST to the base URL plus the endpoint, with an Authorization: Bearer <key> header and a JSON body of model, state and questions. The body carries model for Cloudflare too, even though the endpoint already names the model.
Importing swarms loads a .env file, found by searching upward from the current working directory. Variables already set in your environment are not overridden. You can keep TYPESAFE_API_KEY, CLOUDFLARE_AUTH_TOKEN and CLOUDFLARE_ACCOUNT_ID there.

Constructor

str
default:"jev-latest"
The model that answers the questions. Its name picks the provider’s base URL, endpoint and API key variable.
Optional[str]
default:"None"
The API key. When omitted, it is read from the api_key_env variable. Surrounding whitespace is stripped, and a blank key counts as missing.
Optional[str]
default:"None"
The environment variable that holds the API key. Defaults to the provider’s: TYPESAFE_API_KEY or CLOUDFLARE_AUTH_TOKEN.
Optional[str]
default:"None"
Root URL of the provider’s API. Defaults to the provider’s. A trailing / is removed. When you pass one, no environment placeholders are filled, so a Cloudflare model with an explicit base_url does not need CLOUDFLARE_ACCOUNT_ID.
Optional[str]
default:"None"
Path of the evaluation endpoint, appended to base_url. Defaults to the provider’s. A value you pass is used as is.
float
default:"30.0"
Seconds to wait for each HTTP request.
int
default:"3"
Retries after the first attempt on rate limits, overloads and connection errors. 0 disables retrying. See Retries.
Optional[Dict[str, str]]
default:"None"
Extra HTTP headers sent with every request. They are merged after the default Authorization and Content-Type headers, so they can replace them.
Optional[Dict[str, Any]]
default:"None"
Extra fields merged into every request body after model, state and questions.
Raises ValueError when:
  • The default base URL needs an environment variable that is not set. For a clef model without CLOUDFLARE_ACCOUNT_ID: Set CLOUDFLARE_ACCOUNT_ID in your environment or .env file. This check runs before the API key check.
  • No API key is found: No API key found. Pass api_key or set TYPESAFE_API_KEY in your environment or .env file.
The constructor makes no network calls.

Attributes

str
The model name sent in every request body.
str
The variable the key was read from, or would have been.
str
The resolved root URL, with placeholders filled and no trailing /.
str
The resolved endpoint path.
float
Per-request timeout in seconds.
int
Retries after the first attempt.
Dict[str, Any]
Fields merged into every request body. {} when not set.
Tuple[str, ...]
Class attribute: ("choice", "score", "noul"). build_payload() rejects any other type. Extend it in a subclass to allow more.
The API key and the extra headers are stored in private attributes, so the telemetry that records constructor settings never captures them.

Questions

Pass questions as a dict keyed by ids you choose. The answers come back under the same ids.
Each question has three keys:
str
required
"choice", "score" or "noul".
Any
required
What the model should decide. Usually a string. A structured object is sent unchanged.
Dict | List
Depends on the type. See the table below.
An answer can also carry its type, which parse_response() checks against the question. The shipped examples read score as a position on the 0 to len(criteria) - 1 scale; with three levels, 2.0 is the top. The state is the content the questions are about. It can be a string, a JSON object or a list, and it is sent unchanged. The shipped examples pass an object, such as {"job": ..., "resume": ...}, and name its fields in their instructions.

Methods

run

Ask every question about the state in one request. The questions are validated before anything is sent.
Union[str, Dict[str, Any], List[Any]]
required
Text, JSON object or list the questions are asked about. Cannot be None.
Dict[str, Dict[str, Any]]
required
Question dicts keyed by an id of your choosing. At least one.
Returns the provider’s decoded response, with Cloudflare’s result envelope removed:
str
The model that answered.
Dict[str, Dict[str, Any]]
One answer per question id. Fields depend on the question type.
Dict[str, Any]
The provider’s token usage. The shipped TypeSafe examples read usage["input_tokens"].
Any other fields the provider returns are passed through.

arun

The async form of run(), with the same arguments, validation, retries and return value. Each call opens its own httpx.AsyncClient, so concurrent calls with asyncio.gather are safe, and so is reusing one instance across separate asyncio.run() calls.

choice

Ask a single choice question. Returns the answer dict, with choice, probabilities and confidence.

score

Ask a single score question. criteria lists the levels from lowest to highest. Returns the answer dict, with score, legend, probabilities and confidence.

noul

Ask a single yes/no question. criteria, when given, holds descriptions under the keys "true" and "false" and is sent only when set. Returns the probability, from 0 (no) to 1 (yes), as a float rather than the full answer dict.
choice(), score() and noul() each send one request through run(), with the question under the id "answer". To ask several questions about the same state, put them all in one run() call instead.

list_models

Same as get_decision_models(). It lists the models from every provider, not only this instance’s.

close

Close the pooled HTTP connection used by run() and the single-question helpers. The first sync request opens an httpx.Client and every later sync request reuses it until you call close(). Calling close() twice is safe, and a request after close() opens a new client. arun() does not use the pooled client.

Extension hooks

Override these in a subclass to support a provider with its own wire format. run() and arun() both go through them.

build_headers

Returns the headers for each request. By default: Authorization: Bearer <key>, Content-Type: application/json, then the headers you passed. The key is available to subclasses as self._api_key.

build_payload

Validate the questions and build the JSON request body. By default it returns {"model": ..., "state": ..., "questions": ..., **extra_body}. It raises ValueError for a None state, an empty question dict, an unknown type, or invalid choice or score criteria. An override replaces this validation.

parse_response

Check the decoded response and return it. By default it unwraps a Cloudflare result envelope when answers is not at the top level. It then raises RuntimeError if any question has no answer, or if an answer’s type differs from its question’s. An answer without a type is accepted.

get_decision_models

List the decision models from every provider. Returns the built-in names, merged with each provider’s live list when its credentials are set:
  • TypeSafe: when TYPESAFE_API_KEY is set, GET https://api.typesafe.ai/v1/models.
  • Cloudflare: when CLOUDFLARE_AUTH_TOKEN is set, GET https://api.cloudflare.com/client/v4/accounts/{CLOUDFLARE_ACCOUNT_ID}/ai/models/search?search=clef. The @cf/cloudflare/ prefix is stripped from each name.
Each request sends the key as a bearer token, with a 10 second timeout. Only live names that start with the provider’s prefix (jev or clef) are kept. Names are de-duplicated and keep their order: TypeSafe’s built-in then live names, then Cloudflare’s. The function never raises. A failed request, a malformed body, or a missing CLOUDFLARE_ACCOUNT_ID logs a warning and leaves that provider’s built-in names in place. With no credentials set, it returns the built-ins without a network call:
Results are not cached, so each call with credentials set makes fresh requests.

Retries

run() and arun() retry on these HTTP statuses: 408, 429, 500, 502, 503, 504 and 529. They also retry on httpx.TransportError, which covers connection failures and timeouts. Other error statuses, such as 400, 401 or 422, raise at once.
  • Attempts: up to max_retries + 1 in total.
  • Backoff: min(0.5 * 2 ** attempt, 8.0) seconds, times a random factor between 0.75 and 1.0.
  • Server hints: a retry-after-ms or retry-after header (in seconds) replaces the backoff. A retry-after that is an HTTP date is ignored.
  • Cap: every delay is clamped to between 0 and 60 seconds.
Each retry logs a warning that names the status and the delay.

Errors

Telemetry

When telemetry is on, the constructor records its settings, without the API key or headers. Each run() call, including the ones made by choice(), score() and noul(), emits a DecisionModel.run span that captures the state. arun() does not emit a span. Set SWARMS_TELEMETRY_ON=false to turn telemetry off.

Using it with agents and swarms

A decision model sits next to your agents. You call it with whatever state you have, and branch on the probabilities. These patterns come from the examples in examples/decision_models/.

Route tasks to agents

Use agent descriptions as choice criteria, and add a noul question to reject out-of-scope work:

Screen many items concurrently

arun() opens a client per call, so you can fan out with a semaphore:

More patterns

Guarded group chat

Subclasses GroupChat and scores every reply for hazards before it posts, then passes, flags or blocks it

Decision director

Replaces the LLM director of a HierarchicalSwarm with one that picks the next worker and stops when the answer is complete

Debate referee

Ends a two-agent debate as soon as a round adds no new argument

Ensemble judge

Rates answers from several models on several criteria in one request, and compares cost with an LLM judge

Screening funnel

Screens 500 resumes with arun(), then sends a shortlist to agents

Ticket triage

Asks choice, score and yes/no questions about one ticket in a single request

Agent

The LLM-backed agent you route to, guard or judge

Multi-agent router

LLM-based routing, for comparison

Group chat

The structure the guarded chat example extends

Telemetry

What the DecisionModel.run span records