October 1, 2026
An AI agent asks for the city instead of guessing: Strands Decider checks tool calls
Amazon's model selects from predefined options and scores actions, but does not write code or chat responses.

In the Oct 1 Strands Agents example, Decider checks 2 conditions before requesting the weather. The agent asks the user to specify the city instead of guessing.
Amazon released Strands Decider 2B for individual decisions within an AI agent. Instead of generating a response, the model selects an option or assigns a score. The vendor rules out coding, chatbots and summarization as suitable tasks.
In the example, before `get_weather`, the model checks whether the user has confirmed the arguments and whether the call is premature. For an application, this checks a specific step, such as a tool call.
How to run the model
First request. Install the package with `pip install strands-decider`. The README example routes a payout inquiry to one of the predefined departments:
```bash strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 --state "Help! My payouts have been failing for 3 days!" --choice "Which team should handle this?=billing,sales,retail" ```
On the first run, the package separately downloads the 4.5 GB Qwen/Qwen3.5-2B-Base model: the Decider checkpoint does not include the base weights. The `--device` flag selects `cuda`, `mps` or `cpu`. Without it, the package selects the device automatically.
The CLI supports three question types:
- `choice`; - `noul`; - `score`.
They can be combined in a single request, loading the shared text only once. An HTTP server is available for integration with an application:
```bash strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8000 ```
The server accepts requests through `POST /v1/systemone`. The v19 weights are published on Hugging Face under Apache-2.0.
What results the team reported
Latency in practice. According to the team's measurements as of Oct 1, half of v19 requests on an RTX 3090 finish within 115 ms, and 95% within 299 ms. On an M3 Pro, warmed-up requests shorter than 300 tokens take 153 ms.
In its own JevBench test, the vendor recorded 167 correct answers out of 231, or 72.3%. The model has a context window of 4,096 tokens.
The team provides `training/recipe.sh all` to reproduce training. According to the Oct 1 README, training takes about 11 hours on a single RTX 3090 with 24 GiB of memory, or 1 hour 10 minutes on 8 H100 GPUs.
In the Strands Agents example, a separate model decides whether to call a tool within the agent.
