Skip to main content
A job that runs every thirty seconds for a week will fail sometimes. A rate limit, a dropped connection, a provider hiccup. The question is not whether it fails but what happens next.

The model

A task that raises is logged and retried at the next interval. It does not take the schedule down. CronJob catches the exception inside the job, records it as a failed execution, and returns normally, so the schedule library moves the job to its next run exactly as it does after a success. A failing job runs no more often than a healthy one: with interval="2seconds", a task that always fails is attempted once every two seconds.
Earlier versions got this wrong twice. First, one exception set is_running = False and killed the scheduler thread, while run() returned a job object as though nothing had happened. Later, a raised exception left the job due, so a failing task was retried on every one-second poll instead of waiting out its interval. Both are fixed: failures are recorded and the job waits for its next interval.

Error budgets

Retrying forever is right for transient failures and wrong for a misconfigured job hammering a dead endpoint. max_consecutive_errors draws the line.
  • None (default): never give up. Every failure is logged, and the task runs again at the next interval.
  • An integer: after that many failures in a row, the job stops and run() raises CronJobExecutionError naming the count and the last error.
The counter resets on any success, so a job that fails occasionally never trips the budget. Only a job failing consistently does.
That distinction is the important part: a clean stop() returns, a job that gave up raises. You can tell them apart.

Monitoring while it runs

get_execution_stats() is safe to poll from another thread.

What to watch

error_count on its own is a poor signal: a job running every two seconds for a day will accumulate failures and be perfectly healthy. consecutive_errors is the one that distinguishes noise from breakage.

A complete example

Runs without API keys, since FlakyAgent stands in for an unreliable upstream.
Over sixty seconds the job makes about thirty attempts, one every two seconds whether the last one failed or not. About half succeed and half fail, and it never stops, because the failures almost never stack ten deep in a row.

Budgets across a fleet

With run_many, budgets are per agent. One agent giving up does not stop its siblings:
If the flaky agent exhausts its five, it stops, the stable one keeps running, and run_many raises once blocking ends, naming which job gave up.

Next steps

Multiple Agents

Fleets on mixed cadences

CronJob Quickstart

Start with a single agent

CronJob Reference

Full parameter and method documentation

Runnable Examples

The example files in the repository