Sub-1-Bit LLM Compression via Latent Factorization
github.com/SamsungLabs
[4 comments hidden]
We've been seeing various optimizations towards streaming, that have been much more impactful in the local AI space, e.g. MoE models where the less busy experts are offloaded to slow RAM or even pruned entirely, engram tables that can be read from NVMe instead of sitting around in RAM etc.
Maybe the trick with these extreme quants would be to increase the total parameter count while quanting individual weights, such that maybe the active parameter count comes down, or streaming weights from RAM or disk becomes more efficient, or cache behavior improves? Say, replacing a single 4bit/weight matmul with 3 1bit/weight operations that produce a much closer result than a single 1bit/weight matmul would.
[3 comments hidden]
[14 comments hidden]
It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.
²As to why, I have seen few explanations. But the empirical evidence is there.
[4 comments hidden]
They would be better served with smaller models that can reliably call tools and generate structured outputs without looping or totally hallucinating. These projects exist, but aren't getting amplified.
[3 comments hidden]
Name them
[3 comments hidden]
https://arxiv.org/abs/2603.00042
It is true that block-headed quantization of everything doesn't work. But, as I am sure will read, if you remove the spiky parts of the parameter values, the residue can be dramatically quantized and compressed while retaining performance.
This is a multi-modal sort of compression where you use different techniques for different phenomena. Simply compression all of the weights, each in isolation, ignores the benefits of compressing them collectively.
[hidden]
Yes, there are many way to distribute the error. And LLM are usually not trained to capacity or are prevented from this due to information bottlenecks in the gradiant flow. This allows unlocking wasted capacity by distribtion of the quantization error. But I am pretty certain that there is no way to trick Shannon.
[4 comments hidden]
That being said, there's a slight misconception about models at lower than 4bpw. There's no fundamental reason why a transformer with low bit weights would be inherently incapable of doing high dimensional function approximation, but training a model at one precision and then quantizing to a lower precision means the training loss is never calculated based on the quantized state.
There's a huge difference between "I trained a ternary model from scratch to do X" and "I trained a model at fp16 to do X and then squashed the hell out of it". Quantization Aware Training is the solve, but it's really expensive compared to a one-time, offline translation of existing weights.
[hidden]
Sure, in fact you can train a 1 bit model to approximate any function. But that is not the point. The problem is that the capacity per weight is diminishing below ~4.5bpw.
With QaT, can you compensate for that by adding more weights. So, instead of training a 1B x 4bpw model to capacity saturation you can train a 4B x 1bpw model and achieve the same capacity. But you will end up with the same number of bits in the model, just distributed differently.
Here are some older experiments of mine with very small models: https://github.com/cpldcpu/BitNetMCU/blob/main/docs/document...
I am quite certain that similar behavior is true also for larger models - it may be more difficult to experimentally demonstrate due to information bottlenecks and the flops required to train them to saturation.
[hidden]
Its not so clear to me what their real explanation is. Yes, it is possible to redistribute the quantization error, but this only works when the model is not trained to capacity.
[hidden]
That said, after being satisfied with Poolside.ai’s Laguna XS 2.1 4 bit quant, when I upgraded to a 64G Mac, I now only run the 6 bit quant and the results are better.
Unfortunately, I think that benchmarks are not useful for me, and I need to take time seeing what works for me: this is one reason that when I find a combination of model, harness, etc. that works for me, then I tend to use it for a month or two before looking around for new stuff.
[hidden]
[3 comments hidden]
[hidden]
[hidden]
So, not something anyone would want to run currently, but an indicator that there is still more to squeeze out of lower precisions.
Trellis quantization is a far more approachable enhancement right now, but it doesn't cross the 1-bit barrier (and perhaps doesn't intend to).
augment_me[hidden]
I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance