Overview
A decision model answers typed questions about a piece of content, which it calls the state. You send the state and a set of questions in one request. Each answer comes back as structured data with probabilities, not as generated text.DecisionModel is the client for these models. It supports three question types:
jev-latest is the default model. Set model_name to "clef" or "clef-flash" to use Cloudflare’s Clef models on Workers AI instead.
When to use it
Use a decision model where you would otherwise call an LLM only to make a discrete judgement and then parse its text:- Routing: pick which agent should handle a task, and escalate when confidence is low.
- Guardrails: score every message for hazards before it is posted.
- Early stopping: ask whether a debate or a swarm has stopped making progress.
- Judging: rate many candidate answers against a rubric in one request.
- Screening: run the same questions over hundreds of items concurrently.
Agent when you need generated text, tool calls or multi-step reasoning.
DecisionModel is a standalone HTTP client built on httpx. It is not an Agent, it does not go through LiteLLM, and no swarm structure takes it as a parameter. You call it from your own code, before, between or after agent runs. See Using it with agents and swarms.Import
swarms.structs. The module is swarms.structs.decision_model.
Providers
The model name picks the provider. A name that starts withclef goes to Cloudflare. Every other name goes to TypeSafe, including names that start with jev and gateway ids such as ~typesafe/jev-latest.
Every request is a
POST to the base URL plus the endpoint, with an Authorization: Bearer <key> header and a JSON body of model, state and questions. The body carries model for Cloudflare too, even though the endpoint already names the model.
- TypeSafe
- Cloudflare Clef
- Any other endpoint
Constructor
str
default:"jev-latest"
The model that answers the questions. Its name picks the provider’s base URL, endpoint and API key variable.
Optional[str]
default:"None"
The API key. When omitted, it is read from the
api_key_env variable. Surrounding whitespace is stripped, and a blank key counts as missing.Optional[str]
default:"None"
The environment variable that holds the API key. Defaults to the provider’s:
TYPESAFE_API_KEY or CLOUDFLARE_AUTH_TOKEN.Optional[str]
default:"None"
Root URL of the provider’s API. Defaults to the provider’s. A trailing
/ is removed. When you pass one, no environment placeholders are filled, so a Cloudflare model with an explicit base_url does not need CLOUDFLARE_ACCOUNT_ID.Optional[str]
default:"None"
Path of the evaluation endpoint, appended to
base_url. Defaults to the provider’s. A value you pass is used as is.float
default:"30.0"
Seconds to wait for each HTTP request.
int
default:"3"
Retries after the first attempt on rate limits, overloads and connection errors.
0 disables retrying. See Retries.Optional[Dict[str, str]]
default:"None"
Extra HTTP headers sent with every request. They are merged after the default
Authorization and Content-Type headers, so they can replace them.Optional[Dict[str, Any]]
default:"None"
Extra fields merged into every request body after
model, state and questions.ValueError when:
- The default base URL needs an environment variable that is not set. For a
clefmodel withoutCLOUDFLARE_ACCOUNT_ID:Set CLOUDFLARE_ACCOUNT_ID in your environment or .env file.This check runs before the API key check. - No API key is found:
No API key found. Pass api_key or set TYPESAFE_API_KEY in your environment or .env file.
Attributes
str
The model name sent in every request body.
str
The variable the key was read from, or would have been.
str
The resolved root URL, with placeholders filled and no trailing
/.str
The resolved endpoint path.
float
Per-request timeout in seconds.
int
Retries after the first attempt.
Dict[str, Any]
Fields merged into every request body.
{} when not set.Tuple[str, ...]
Class attribute:
("choice", "score", "noul"). build_payload() rejects any other type. Extend it in a subclass to allow more.Questions
Pass questions as a dict keyed by ids you choose. The answers come back under the same ids.str
required
"choice", "score" or "noul".Any
required
What the model should decide. Usually a string. A structured object is sent unchanged.
Dict | List
Depends on the type. See the table below.
An answer can also carry its
type, which parse_response() checks against the question. The shipped examples read score as a position on the 0 to len(criteria) - 1 scale; with three levels, 2.0 is the top.
The state is the content the questions are about. It can be a string, a JSON object or a list, and it is sent unchanged. The shipped examples pass an object, such as {"job": ..., "resume": ...}, and name its fields in their instructions.
Methods
run
Union[str, Dict[str, Any], List[Any]]
required
Text, JSON object or list the questions are asked about. Cannot be
None.Dict[str, Dict[str, Any]]
required
Question dicts keyed by an id of your choosing. At least one.
result envelope removed:
str
The model that answered.
Dict[str, Dict[str, Any]]
One answer per question id. Fields depend on the question type.
Dict[str, Any]
The provider’s token usage. The shipped TypeSafe examples read
usage["input_tokens"].arun
run(), with the same arguments, validation, retries and return value. Each call opens its own httpx.AsyncClient, so concurrent calls with asyncio.gather are safe, and so is reusing one instance across separate asyncio.run() calls.
choice
choice question. Returns the answer dict, with choice, probabilities and confidence.
score
score question. criteria lists the levels from lowest to highest. Returns the answer dict, with score, legend, probabilities and confidence.
noul
criteria, when given, holds descriptions under the keys "true" and "false" and is sent only when set. Returns the probability, from 0 (no) to 1 (yes), as a float rather than the full answer dict.
choice(), score() and noul() each send one request through run(), with the question under the id "answer". To ask several questions about the same state, put them all in one run() call instead.
list_models
get_decision_models(). It lists the models from every provider, not only this instance’s.
close
run() and the single-question helpers. The first sync request opens an httpx.Client and every later sync request reuses it until you call close(). Calling close() twice is safe, and a request after close() opens a new client. arun() does not use the pooled client.
Extension hooks
Override these in a subclass to support a provider with its own wire format.run() and arun() both go through them.
build_headers
Authorization: Bearer <key>, Content-Type: application/json, then the headers you passed. The key is available to subclasses as self._api_key.
build_payload
{"model": ..., "state": ..., "questions": ..., **extra_body}. It raises ValueError for a None state, an empty question dict, an unknown type, or invalid choice or score criteria. An override replaces this validation.
parse_response
result envelope when answers is not at the top level. It then raises RuntimeError if any question has no answer, or if an answer’s type differs from its question’s. An answer without a type is accepted.
get_decision_models
- TypeSafe: when
TYPESAFE_API_KEYis set,GET https://api.typesafe.ai/v1/models. - Cloudflare: when
CLOUDFLARE_AUTH_TOKENis set,GET https://api.cloudflare.com/client/v4/accounts/{CLOUDFLARE_ACCOUNT_ID}/ai/models/search?search=clef. The@cf/cloudflare/prefix is stripped from each name.
jev or clef) are kept. Names are de-duplicated and keep their order: TypeSafe’s built-in then live names, then Cloudflare’s.
The function never raises. A failed request, a malformed body, or a missing CLOUDFLARE_ACCOUNT_ID logs a warning and leaves that provider’s built-in names in place. With no credentials set, it returns the built-ins without a network call:
Retries
run() and arun() retry on these HTTP statuses: 408, 429, 500, 502, 503, 504 and 529. They also retry on httpx.TransportError, which covers connection failures and timeouts. Other error statuses, such as 400, 401 or 422, raise at once.
- Attempts: up to
max_retries + 1in total. - Backoff:
min(0.5 * 2 ** attempt, 8.0)seconds, times a random factor between 0.75 and 1.0. - Server hints: a
retry-after-msorretry-afterheader (in seconds) replaces the backoff. Aretry-afterthat is an HTTP date is ignored. - Cap: every delay is clamped to between 0 and 60 seconds.
Errors
Telemetry
When telemetry is on, the constructor records its settings, without the API key or headers. Eachrun() call, including the ones made by choice(), score() and noul(), emits a DecisionModel.run span that captures the state. arun() does not emit a span. Set SWARMS_TELEMETRY_ON=false to turn telemetry off.
Using it with agents and swarms
A decision model sits next to your agents. You call it with whatever state you have, and branch on the probabilities. These patterns come from the examples inexamples/decision_models/.
Route tasks to agents
Use agent descriptions aschoice criteria, and add a noul question to reject out-of-scope work:
Screen many items concurrently
arun() opens a client per call, so you can fan out with a semaphore:
More patterns
Guarded group chat
Subclasses
GroupChat and scores every reply for hazards before it posts, then passes, flags or blocks itDecision director
Replaces the LLM director of a
HierarchicalSwarm with one that picks the next worker and stops when the answer is completeDebate referee
Ends a two-agent debate as soon as a round adds no new argument
Ensemble judge
Rates answers from several models on several criteria in one request, and compares cost with an LLM judge
Screening funnel
Screens 500 resumes with
arun(), then sends a shortlist to agentsTicket triage
Asks choice, score and yes/no questions about one ticket in a single request
Related
Agent
The LLM-backed agent you route to, guard or judge
Multi-agent router
LLM-based routing, for comparison
Group chat
The structure the guarded chat example extends
Telemetry
What the
DecisionModel.run span records