Strands Decider 2B: a small, open-source, decision model
strandsagents.com
[8 comments hidden]
Meet Jerry- it's literally a guy named Jerry answering your questions.
[3 comments hidden]
Jerry, "The Decider" https://www.youtube.com/watch?v=r8VbzrZ9yHQ
[hidden]
[12 comments hidden]
[6 comments hidden]
[4 comments hidden]
[3 comments hidden]
[0] https://news.ycombinator.com/item?id=49723267 (see parent for reference)
[2 comments hidden]
So "noul"--short for Bernoulli as in the Bernoulli distribution--"maps to if-statements."
[hidden]
I'd googled "bernoulli map" after reading that, and came across this[0], which I wasn't aware of and thought was somehow related so I entrenched my misunderstanding (I didn't dig in.)
Side note: I just wrote the sentence "you're completely right!", and almost changed it because it sounds like slop now. I wonder if we're gonna get an increase in these sentences in human written prose over time as people mimic these sentences, or a decrease as people shy away from them to not sound like AI. :)
[hidden]
[hidden]
[hidden]
[hidden]
It is from BerNOULli distribution [0]
[0] https://en.wikipedia.org/wiki/Bernoulli_distribution
EDIT: formatting
[5 comments hidden]
It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?
[2 comments hidden]
[hidden]
On humans-for-humans, blog posts with my name on aren't written by AI (https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html) either for my personal blog or at work. I'm a heavy AI user, but this is something I think is best left to humans.
There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence.
We're releasing more of that functionality in strands, and I expect an explosion of innovation in that area.
[2 comments hidden]
It is just the right mix of size, capability and speed to make it generally useful for adhoc bulk classification tasks.
For those wanting to run it in browser: https://huggingface.co/alxnahas/strands-decider-2B-webgpu
[hidden]
[2 comments hidden]
Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:
$ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose
[...]
File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name
raise ValueError(f"Can not map tensor {name!r}")
ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'[hidden]
I haven't looked in depth, but it should be doable without modifications to llama.cpp. You can't just convert the LoRA, through - there's a whole pointer head and some custom layers that need to be correctly handled (in fact, the LoRA mostly exists to get the base model to behave the way the pointer head needs).
[hidden]
Even though they weren't themselves decision models exactly, they were so fast that you could use them in a similar way. They obliterated Qwen and other models for a lot of tasks in that use case.
I'm not sure I would agree with some of the claims made in this article though. You can definitely do a lot of similar tasks to regular LLMs with decision models, you just have to chain the results. You can even technically infuse a broader perspective into each token choice or force certain context to have priority in the decision making of the next token.
[6 comments hidden]
[2 comments hidden]
If I have a sizeable amount of labeled data and need decisions calibrated to that data should I
1. Ignore the hype and train a traditional classifier
2. Finetune an LLM based decision model
3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow
If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?
[hidden]
(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.
(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.
(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.
The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.
[hidden]
The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results
[2 comments hidden]
Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.
As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).
Everything you need to fine-tune, or even retrain from scratch, is in the github repo.
[2 comments hidden]
[2 comments hidden]
[hidden]
I appreciate what the JevBench folks are doing, but it's true the JevBench (like all benchmarks) isn't representative of every real-world workloads. I particularly don't love how JevBench handles confidence, and rewards being highly confident in certain cases. Benchmarking is hard, and JevBench is probably more representation of its target workloads than TPC-C is :)
As for actual intelligence, there's only so much you can expect from 2B without reasoning. One of the design challenges in training this model was to avoid forgetting too much, and the KL-to-frozen-base step partially exists for that purpose. Even then, Qwen3.5 2B base has fairly limited single-pass reasoning ability (which shows up for us in the performance on the JevBench hard set).
[7 comments hidden]
[hidden]
(Disclaimer, I work at Cloudflare, but not on models)
[4 comments hidden]
[3 comments hidden]
I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking.
Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.
[2 comments hidden]
[hidden]
v21 supports vision: https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...
[hidden]
[4 comments hidden]
[hidden]
{ "model": "strands-decider-2B-hobson-v19", "answers": { "is_urgent": { "type": "noul", "noul": 0.8287 } }, "usage": { "input_tokens": 86, "output_tokens": 1 }, "latency_ms": 1732.17 }
This is how I got it running - https://gist.github.com/2891eb0db9ea92c1a4e860d44f556292
There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.
[5 comments hidden]
[hidden]
To vaguely confirm that this is plausible, i tested the sample query from the post with differently arranged options (This should produce slightly different outputs since each option influences all subsequent hidden states, including those of the other options):
"billing,sales,retail" produces billing -> 0.843, retail -> 0.092, sales -> 0.065
"retail,billing,sales" produces billing -> 0.470, retail -> 0.468, sales -> 0.062
"retail,sales,billing" produces billing -> 0.517, retail -> 0.415, sales -> 0.068
"billing,retail,sales" produces billing -> 0.803, retail -> 0.146, sales -> 0.051
"sales,retail,billing" produces billing -> 0.647, retail -> 0.127, sales -> 0.225
So each of these arrangements favor billing, but with pretty different confidence scores, and it looks like there is a pretty big bias toward the first option listed - which makes me wonder how viable it would be to have the torso process these options independently instead
Their docs say:
> An option's score depends on what it says, not where it sits. There is no per-option parameter anywhere in the model, so nothing can learn that "the first option is usually the right one"
This isn't quite correct. There might not be an explicit "per-option" parameter, but the hidden states themselves "contain" all prior options. My gut feeling tells me they messed up data augmentation by incorrectly permuting options during training.
Edit: the above numbers were obtained with the v19 version/checkpoint, v21 is much less sensitive to ordering, but still shows a first position bias.
[0] https://github.com/strands-labs/strands-decider/blob/main/do...
[3 comments hidden]
On multiple tokens, we put each option (however many tokens it is) onto a line, and then read the hidden state at the end of that line to use at the option state. This is done option-by-option.
The hidden state at the end of the question is the query `q` (again, not sensitive to how many tokens the question is), each options end-of-line state is the key `k`, and the per-option logit is calculated as `q.k / sqrt(256)`.
[2 comments hidden]
[hidden]
I should let the development agent select a model and run it to improve token efficiency.
Is it being used this much these days?
real_faxenoffAs a regular user of a bunch of specialized micromodels, I'll tell you this: you won't be happy with such a model (and its JEV counterparts)[10 comments hidden]
NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).
I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.
cyanydeezI haven't seen any frameworks for running the NPU. my 395+ needs a buddy.[5 comments hidden]
AbsurdCensorI think under Lemonade you can do that, especially with AMD systems, using the models that allow for hybrid operation. Prefill happens on th[2 comments hidden]
shifto[hidden]
InfernalFastFlowLM is what you're looking for - Lemonade does package it, as sibling comment says. https://fastflowlm.com/[2 comments hidden]
https://fastflowlm.com/
p_l[hidden]
ricardobeatMost LLMs cannot run efficiently on current NPUs (except for prefill stage), the hardware was built for a different kind of ML workload.[3 comments hidden]
mermericoDecision models are prefill only[2 comments hidden]
spwa4[hidden]
nicoWhat models are you running on CPU? Any repos or gists you can share? Curious about the models and your use case. Are you doing multi-langua[hidden]
Btw, would love your opinion on this: pre trained classifiers that run and train on CPU https://github.com/nicobrenner/jeffy