Skip to content

Technical data sheet

Models & hardware

Your model is created by fine-tuning an open base model: Qwen3.5-9B for size M, Qwen3.8-27B for size L. Fine-tuning runs in 8-bit precision (FP8); the model ships in NVFP4 (4-bit) for Blackwell cards and in BF16 for all other supported cards.

M

Balanced standard for most departments: broader knowledge base, more nuanced answers.

Base model
Qwen3.5-9B
Model size
9B
Data volume
ZIP upload up to 8 GB
What it can realistically do

The 9B model is the recommendation for one coherent area of the business with consistent language: one product area, one department, one set of rules. It summarises longer documents, writes emails, quotes and internal texts in your company's style, and answers questions across a larger knowledge base, including nuanced and ambiguous questions. It distinguishes stable company knowledge from day-to-day figures: it does not assert prices, stock levels or contact details; it points to the reliable source. Limits: for very large knowledge bases or particularly analytical tasks, the 27B is the right choice.

L

The largest model, for large knowledge bases and demanding questions.

Base model
Qwen3.8-27B
Model size
27B
Data volume
ZIP upload up to 16 GB
What it can realistically do

The 27B model is intended for companies that consolidate a lot of knowledge and need high answer quality. It works across very large document sets and keeps track over very long conversations and documents thanks to a native context of 262,000 tokens; it covers 201 languages. It suits tasks where an answer has to connect several threads. Limits: it is focused on your knowledge, not a universal model like the large cloud services; in return it runs entirely locally and belongs to you. Its hardware requirement is the highest.

Applies to both models: fine-tuned in FP8, shipped for your hardware in NVFP4 (4-bit) or BF16, and built on an Apache 2.0 base with full transfer of rights including a certificate of ownership.

Recommended hardware

You source the hardware yourself. Availability and prices on the market currently fluctuate considerably; rather than tie you to a specific device or price, we leave the choice to you. The following NVIDIA systems serve as a guide; our setup program checks your hardware and selects the right path for your model.

Recommended

NVIDIA DGX Spark

GB10 Grace Blackwell Superchip

Our recommendation for most: compact as a book, quiet, sits on a shelf and carries your fine-tuned model with headroom.

AI performance
up to 1 petaFLOP (FP4)
Memory
128 GB unified (CPU + GPU, LPDDR5X)
Processor
20 Arm cores (GB10)
Concurrent users
small teams, approx. 5–15 active chats smoothly
Storage
4 TB NVMe
Networking
10 GbE, Wi-Fi 7, ConnectX-7 (200 Gb/s)
Multi-user
one shared instance
Form factor
150 × 150 × 50.5 mm, approx. 1.1 litres
More compute

NVIDIA DGX Station

GB300 Grace Blackwell Ultra Superchip

For substantially more compute: several teams, larger models, one device. Data-centre class in a tower under the desk.

AI performance
up to 20 petaFLOPS (FP4)
Memory
748 GB coherent (incl. HBM3e GPU memory)
Processor
72-core NVIDIA Grace
Concurrent users
entire departments, 100+ active chats smoothly
Storage
NVMe SSD, configuration-dependent
Networking
ConnectX-8, up to 800 Gb/s
Multi-user
up to 7 isolated instances (MIG)
Form factor
tower, liquid-cooled

User figures: concurrently active chats at our model sizes (9B–27B), conservatively estimated. The systems serve as a guide; you source the hardware yourself.

Recommended hardware for local operation · IonKon