Research
The physical maximum. Locally.
Local AI is the right answer to the data question: your documents, your queries and your model stay on your premises. But local AI has to be efficient. Large cloud providers compensate for inefficient models with massive compute. A company with a server in a cabinet cannot.
We tailor AI architectures to customers and work at the same time on making them fundamentally more efficient: training and inference methods that make AI leaner and more sustainable without giving up capability. Our goal is to offer you the physical maximum in the long run: the best answer quality that can be extracted from every watt and every byte of memory bandwidth.
Precision ladder
16Bit
BF16 · Industry standard
The usual path: high precision, high memory and energy demand.
4Bit
NVFP4 · In production today
Our customer models: fine-tuned in FP8, shipped in 4 bit, ready to run on your hardware.
1Bit
1 bit · Our research
Our lab: trainable 1-bit architectures for standard hardware.
Precision ladder · bits per weight
Where we stand today
We fine-tune customer models in 8-bit precision (FP8) and ship them in NVFP4 (4-bit) for Blackwell cards, in BF16 for all other cards. Going from 8 bit in training to 4 bit in operation is everyday production for us and the basis of our research.
What this means for you
More model per device: the same hardware carries a larger model or answers faster. In the longer term, smaller, cheaper and more frugal devices become possible, and the energy per answer drops.
1-bit architectures
Our current focus is on extremely low-precision model architectures, down to one bit per weight. Such models need only a fraction of the memory and compute of today's models, and therefore run fast even on standard hardware, without special accelerators or a data centre.
The challenge lies in training such models, less in running them. For this we have developed our own architectures and training methods; we deliberately do not publish their details.
In our experiments, this yields substantial speed-ups on standard hardware without a notable loss in answer quality.
Working on efficient AI yourself?
We are happy to talk about collaborations, joint experiments or our methods.
Get in touch