Skip to main content
A counting problem goes wrong in one of a few places: the modular reduction, a missed residue, or an off-by-one at a boundary. A breadth-first search keeps three lines of work alive at each level, and averaging two ratings per step stops one over-optimistic rating from steering the beam.

Prerequisites

TreeOfThoughts is not in swarms 15.0.3 on PyPI, so install from GitHub until the next release.

The code

number_theory_counting.py
It prints the answer, the steps on the best path, and what the search cost.

What the settings do

  • breadth=3 keeps the three best open paths at each level instead of the default two.
  • n_evaluate_samples=2 rates every candidate twice and averages the two scores.
  • max_depth=4 leaves room for the usual path: reduce modulo 7, test the residues, count each residue class, add.
  • thought_description sizes a step as one deduction with its computation shown, so the evaluator can check it.
  • evaluation_criteria names the three ways this problem goes wrong: bad modular arithmetic, a missed residue, and an off-by-one at 1 or 1000.

Check the answer

Test the seven residues of n modulo 7: Only n ≡ 2 and n ≡ 4 (mod 7) work. From 1 to 1000 there are 143 of each (2, 9, …, 996 and 4, 11, …, 998), so the answer is 286.

Cost

This configuration expands at most 10 nodes and makes at most 71 model calls: 10 generation calls, 60 ratings (3 candidates × 2 ratings per expanded node) and 1 answer call. See Cost for the formula.

Next

Fermi estimate

Sampled steps compared by vote.

TreeOfThoughts reference

Every parameter, strategy, and result field.