An LLM with the language taken out, so it just decides
Strands Decider strips a 2B model's ability to generate text and replaces it with a pointer head that picks options and rates things — with calibrated confidence, in ~115 ms.
What it does
Strands Decider is a “decision model”: it takes some state (text, or images with --vision) and answers questions — pick one of N options, yes/no, or score on an ordered rubric. It’s aimed at the small routing decisions inside agentic workflows: which LLM should handle this, which team owns this ticket, is this urgent. The reference model is strands-decider-2B-hobson-v21, 1.9B parameters, served via CLI or an HTTP endpoint.
The interesting bit
The architecture is the story: take Qwen3.5-2B-Base, cut off its language-modelling head, and bolt on a ~1M-parameter pointer head that scores each option by comparing hidden states. No generation loop, one forward pass. Because the head has no per-option parameters, label sets are defined by the request, not baked into the weights — nothing can learn “the first option is usually right,” and there’s no cap on option count.
Key highlights
- Median 115 ms per question on an RTX 3090; also serves on Apple silicon (faster still via MLX) and CPU.
- Calibrated confidence on every answer: at 0.9+ confidence on unseen short classification tasks, it’s right ~95% of the time — something the README notes frontier LLM APIs don’t expose.
- Many questions over one text are nearly free: the text is read once, each question adds only its own tokens.
- Vision works with no image training — v19 matched an image-trained 2B decider and was better calibrated on NaturalBench (ECE 0.014 vs 0.080).
- JevBench v1 public set: 0.762 accuracy (176 of 231 tasks); 1.000 on the easy tier, 0.550 on hard.
Caveats
- The MLX extra “ships with the next release” — until then you need to install from a clone.
- The README itself flags that 231 tasks is few: differences under ~10 tasks between single runs are unresolved noise, and the v21-vs-v19 comparison should be read as two single runs, not a measured gain.
- Hard-tier accuracy is 0.550 — this is a fast decider, not a reasoner.
Verdict
If you’re building agents and burning LLM tokens on routing and triage decisions, this is worth a look — fast, cheap, and honest about its own uncertainty. If you need generated text or hard reasoning, it’s the wrong tool by design.
Frequently asked
- What is strands-labs/strands-decider?
- Strands Decider strips a 2B model's ability to generate text and replaces it with a pointer head that picks options and rates things — with calibrated confidence, in ~115 ms.
- Is strands-decider open source?
- Yes — strands-labs/strands-decider is open source, released under the Apache-2.0 license.
- What language is strands-decider written in?
- strands-labs/strands-decider is primarily written in Python.
- How popular is strands-decider?
- strands-labs/strands-decider has 504 stars on GitHub.
- Where can I find strands-decider?
- strands-labs/strands-decider is on GitHub at https://github.com/strands-labs/strands-decider.