Skip to content

Research

The physical maximum. Locally.

Local AI is the right answer to the data question: your documents, your queries and your model stay on your premises. But local AI has to be efficient. Large cloud providers compensate for inefficient models with massive compute. A company with a server in a cabinet cannot.

We tailor AI architectures to customers and work at the same time on making them fundamentally more efficient: training and inference methods that make AI leaner and more sustainable without giving up capability. Our goal is to offer you the physical maximum in the long run: the best answer quality that can be extracted from every watt and every byte of memory bandwidth.

Precision ladder

16Bit

BF16 · Industry standard

The usual path: high precision, high memory and energy demand.

4Bit

NVFP4 · In production today

Our customer models: fine-tuned in FP8, shipped in 4 bit, ready to run on your hardware.

1Bit

1 bit · Our research

Our lab: trainable 1-bit architectures for standard hardware.

Precision ladder · bits per weight

Where we stand today

We fine-tune customer models in 8-bit precision (FP8) and ship them in NVFP4 (4-bit) for Blackwell cards, in BF16 for all other cards. Going from 8 bit in training to 4 bit in operation is everyday production for us and the basis of our research.

What this means for you

More model per device: the same hardware carries a larger model or answers faster. In the longer term, smaller, cheaper and more frugal devices become possible, and the energy per answer drops.

1-bit architectures

Our current focus is on extremely low-precision model architectures, down to one bit per weight. Such models need only a fraction of the memory and compute of today's models, and therefore run fast even on standard hardware, without special accelerators or a data centre.

The challenge lies in training such models, less in running them. For this we have developed our own architectures and training methods; we deliberately do not publish their details.

In our experiments, this yields substantial speed-ups on standard hardware without a notable loss in answer quality.

Working on efficient AI yourself?

We are happy to talk about collaborations, joint experiments or our methods.

Get in touch
Research: 1-bit models and local limits · IonKon