Microsoft-Decision-1, our model for fast decision-making
commandline.microsoft.com
[12 comments hidden]
[3 comments hidden]
[2 comments hidden]
[hidden]
https://blogs.windows.com/windowsexperience/2026/10/07/build... https://github.com/microsoft/windowsML
[2 comments hidden]
[hidden]
[hidden]
[3 comments hidden]
Supports GPU, NPU and CPU.
[2 comments hidden]
[hidden]
[hidden]
[hidden]
Qwen tunes are nice and all but either price or performance makes each example I have tested not viable unless you are able to cheaply self-host and fine-tune further.
[2 comments hidden]
Using a decoder model for this comes down to picking the most probable class based on the probability what the next token assigned to a class is. They are all probabilities but semantically mean totally different things. Am I right?
[2 comments hidden]
[5 comments hidden]
[2 comments hidden]
Also, section 2.3: https://typesafe.ai/legal/mca
[hidden]
It does say "develop (or to facilitate the development of) a similar or competing product or service", but I think it would be a long stretch to say that's the case if they would just publish benchmarks. Microsoft legal department might disagree.
[3 comments hidden]
Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna).
I can't remember when a new product created So many competitors so quickly. What is very clear is that everyone is saying "Doh!", slapping themselves on the forehead, and scrambling to get a slice of this obvious-in-retrospect massive pie.
What no-one appears to have done yet is to come close to Jev on pricing!
[hidden]
[2 comments hidden]
I guess the fear of Chinese models is finally subsiding.
[4 comments hidden]
[hidden]
[2 comments hidden]
Question: “Which team should handle this message?”
Result: “Tech Support (85%)”
Yep, sounds about right.
[hidden]
[hidden]
[hidden]
[2 comments hidden]
[hidden]
https://www.whitehouse.gov/presidential-actions/2025/01/endi...
No need to use AI, this was done by ctrl-f on dangerous terms such as "disabilities", "trauma" and "female".
(No, I am not making this up)
nejch[19 comments hidden]
Kind of funny how much hype they can all get out of this, but Qwen really is the little engine that could. Great to see open weights (if not open source) driving the whole ecosystem like this though.
manmal[6 comments hidden]
girvo[5 comments hidden]
Super cool finding!
ByteAtATime[4 comments hidden]
RussianCow[hidden]
anewhnaccount2[2 comments hidden]
manmal[hidden]
NitpickLawyer[6 comments hidden]
The insistence of naming it open weights as opposed to open source is getting ridiculous, and it's both irrelevant (i.e. no one cares in practice) and factually incorrect.
Weights are source in language models. Apache defines source as ""Source" form shall mean the preferred form for making modifications". That is precisely what's happening here. Everyone is using the preferred form for making modifications to these models (including the model creators themselves). A model is "created" at init time, and then "trained" by modifying the weights.
All these models are open source. What's not open sourced (with qwen et all) is the training code. So open source model, no training code. And that's ok. There are labs that release those as well. Apertus and Olmo series come with open source models, open source training and open datasets. Nemotron comes with open source models, open source training and some open datasets, while others are not published. And that's ok too.
The fact that you see all these models being modified (from AR completion models to "decision models") and re-released should be all the proof you need. That's what a license offers you. The right to inspect, run, modify and re-release a model. A license cannot (and never did) give you any other rights. OpEnWeIgHtS is silly.
nejch[2 comments hidden]
I'm happy to have and be able to serve these models and see the ecosystem thrive. And lots of open innovation is outside of weights anyway as DeepSeek repeatedly shows.
mh-[hidden]
Open source has historically meant you could download the source and build your own binary. The appropriate analogue here, IMHO, to the build->binary process is training->weights.
Rohansi[hidden]
Dictionary.com defines source as:
> any thing or place from which something comes, arises, or is obtained; origin.
teruakohatu[2 comments hidden]
Weights are source in the same way as any x86 binary is source.
You easily modify a x86 binary and change behaviour or examine the machine code instructions. You probably are not aware how easy it is to change the behaviour of a binary executable.
gunalx[hidden]
Also if compiling literary costs millions of dollars, I would prefer the precompiled one.
giancarlostoro[2 comments hidden]
throwa356262[hidden]
And also, it is nowhere as good as Qwen
sidd0103[hidden]
mjb[3 comments hidden]
As of this afternoon, we also have Gemma4-based variants of strands-decider at 2B, 4B, 12B, and 26B: https://huggingface.co/StrandsAgents
I don't know if Fabio (who's leading work on this model line) would agree, but if I was going to start over I'd probably pick Gemma4 as the base rather than Qwen3.5. But it's not surprising to see a lot of Qwen competition, and I totally agree it's nice to see the work open.
nonatofabio[hidden]
nowittyusername[hidden]