Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

Colibri - Inference

Colibri

Colibri streams MoE experts from NVMe so frontier models never have to fit in RAM. Kimi K3 on 32 GB, GLM-5.2 on 16 GB, pure C, Apache 2.0. And about 1.8 tokens a second.

License Apache 2.0
License Apache 2.0
TL;DR
  • Dense layers in RAM, routed experts streamed from NVMe on demand
  • Runs 744B GLM-5.2 on 16 GB and 2.8T Kimi K3 on 32 GB of RAM
  • Pure C, zero engine dependencies, and roughly 1 to 2 tokens per second
System Requirements
RAM8 GB for OLMoE, 16 GB for GLM-5.2, 32 GB+ for Kimi K3
GPUNone required; GPU optional via CUDA, Metal or Vulkan
VRAMNone. Disk is the bottleneck, 372 GB for GLM-5.2, 1.6 TB for Kimi K3
✓ Apple Silicon

This is part four of our local inference performance series, and it takes the argument somewhere the first three did not go. Part one shrank the model. Part two made decoding faster without changing the output. Part three spent the bits more carefully. Colibri asks a different question: what if the model never has to fit at all? It runs Kimi K3, all 2.8 trillion parameters of it, on a desktop with 32 GB of RAM, by leaving the experts on your SSD and fetching them as the router asks for them. It is pure C, Apache 2.0, and it has collected 38,000 stars since July.

Storage, RAM and VRAM as one hierarchy

The usual rule for running a Mixture-of-Experts model locally is that the weights have to fit in memory, so a 744B model is simply out of reach unless you own a server. Colibri rejects the premise. It treats storage, RAM and VRAM as a single inference hierarchy and moves weights between them on demand, which the project describes as being like a just-in-time compiler, except for weights instead of code.

The split is what makes it work. In an MoE model only a small fraction of the parameters fire for any given token. For GLM-5.2 that is 40B active out of 744B total. Colibri keeps the dense layers resident, about 9.9 GB at int4, and leaves the 19,456 routed experts on NVMe at roughly 19 MB each, 372 GB in total. When the router picks experts for a layer, those get read from disk.

Four tricks stop that from being unusably slow. There is a per-layer LRU cache with learned pinning, so experts you keep hitting stay in fast memory. There is one-layer-ahead prefetch, which works because the project measured routing to be 71.6% predictable one layer ahead. There is batch-union reading, so an expert needed twice at the same token position is fetched once. And the engine writes usage files as it runs, then pins the hot experts to faster tiers automatically on later runs.

What it will run, and on what

The supported list is essentially our Models back catalogue, which is a useful coincidence for anyone who has been reading along.

ModelTotal / activeDiskRAM needed
OLMoE7B / 1B~7 GB8 GB
Qwen3.6-35B-A3B35B / 3B~20 GB24 GB
DeepSeek V4 Flash284B / 13B167 GB16 GB minimum
DeepSeek V4.1 Flash552B / 16B203 GB16 GB minimum
GLM-5.2744B / 40B372 GB16 GB, 24 comfortable
Inkling975B / 41B469 GB25 GB
Kimi K32.8T / 104B~1.6 TB32 GB+

No GPU is required. A GPU only buys speed. Kimi K3 and DeepSeek V4 load in their native MXFP4 and fp4 formats with no conversion step, and Qwen3.8-Flash-Next uses its native block-FP8. Platforms covered are Linux on x86-64 and aarch64, macOS on Apple Silicon through Metal, Windows 11 natively, and WSL2, with optional CUDA, Metal or Vulkan backends. The Vulkan path reaches AMD cards that ROCm has dropped, including an RX 580.

Now the number you actually need

Everything above is the good news. Here is the trade, in the project's author-reported figures.

SetupDecode speed
25 GB dev box, cold cache0.05 to 0.1 tok/s
Single RTX 5070 Ti1.07 tok/s
128 GB CPU-only desktop, warm~1.8 tok/s
Six RTX 5090s, full residency5.8 to 6.8 tok/s

On a cold 25 GB machine you are looking at a token every ten to twenty seconds. A 200-token reply is most of an hour. This is not a chat experience, and the project does not pretend otherwise. Its own framing is blunt: "Speed is set by your disk, because the experts are streamed from it."

What you are buying is not throughput, it is access. A 744B model answering slowly on hardware you already own is a different proposition from a 744B model you cannot run at all. For batch work, overnight jobs, evaluation runs or anything where latency does not matter, that is a real capability. For anything interactive, it is not.

Your disk matters more than your GPU. Two independent NVMe drives measured 37.5% faster decode than one, and enabling DIRECT I/O on Windows with a Blackwell card gave another 34%, taking raw throughput from 4.25 to 9.69 GB/s. Both gains are hardware-dependent and the project says so: shared controllers, DRAM-less or QLC drives, and virtualised disks can make either change neutral or negative.

Running it

The engine has no dependencies. No BLAS, no Python at inference time, just a C compiler with OpenMP. Python appears once, for the one-time model conversion.

git clone https://github.com/JustVugg/colibri && cd colibri/c
./setup.sh
make -C c glm

Then convert a model and point the engine at it. Conversion downloads and converts shard by shard, so you do not need the whole thing on disk twice.

./coli convert --model /nvme/glm52_i4

COLI_MODEL=/nvme/glm52_i4 ./coli doctor
COLI_MODEL=/nvme/glm52_i4 ./coli tune
COLI_MODEL=/nvme/glm52_i4 ./coli chat

Run doctor before anything else, it is a readiness check, and tune measures the fastest profile for your specific disk rather than guessing. There is also plan to inspect where weights will land, serve for a headless API and web for an API with a dashboard. If you have two drives, the mirror planner will split the model across both with a budget.

The parts worth knowing about

Two features stand out beyond the streaming itself. Brio mode scores probabilities read-only without generating anything, measured at 2.4 times faster on a four-field schema and 5.7 times on four questions, which is a sensible answer for classification and routing work where you do not need prose. And the KV state persists to disk compressed, 57 times smaller for MLA models at 576 floats per token against 32,768, so a conversation can restart warm and, the project claims, byte-identical to an uninterrupted session.

The engineering culture on display is unusually careful. The forward pass is validated against a teacher-forced reference implementation, typically matching 30 to 32 positions out of 32. Sparse attention is verified by forcing full-key selection and checking it reproduces dense attention exactly. And the project publishes the negative results too: speculative decoding on DeepSeek V4 was measured at acceptance rates of 1 in 15 and 10 in 24, judged not to repay the verification cost, and shipped disabled by default. Quantizing the MTP head to int4 dropped acceptance to between zero and four percent, so int8 is required there.

That is a rare thing to find in a README. The stated principle is "no SLA on speed, and a hard guarantee on semantics", which is exactly the right way round for something this aggressive: insufficient fast memory may make it slow, but it must not quietly change what the model says.

Where it fits against the rest of the series

The four approaches are not competitors so much as different answers to "my hardware is too small".

ApproachWhat it changesWhat it costs
Ternary trainingModel is smaller from the startKnowledge and vision accuracy, custom runtime
Smart quantizationFewer bits where they matter leastAlmost nothing at 3.5 bits, runs in stock llama.cpp
Exact draftingMore tokens per second, same outputComplexity, a pinned MLX version
Colibri streamingModel does not need to fitSpeed, and a lot of it

If a quantized 27B does what you need, parts one and three are the cheaper answer and you will get interactive speeds. Colibri is for when the model you want is a frontier MoE and there is no quantization that makes it fit. That is a narrower case, but it used to be an impossible one.

Who should run this

Run it if you want to evaluate a frontier open model on your own hardware without renting a server, if your workload is batch or overnight, or if you are doing research where a slow honest answer beats no answer. Run it if you have fast NVMe and are willing to let the tuner find your profile.

Do not run it expecting a local chatbot. At one to two tokens per second on a good desktop, and a tenth of that on a cold small one, the interaction model is closer to submitting a job than holding a conversation. And keep an eye on the disk cost: 372 GB for GLM-5.2 and 1.6 TB for Kimi K3 are real storage commitments before you generate a single token.

Sources and further reading

Ten minutes: clone it, build with the setup script, and run coli doctor and coli plan against a model directory before you download 372 GB of anything. The planner will tell you exactly where the weights would land on your machine and what it expects your disk to give you, which is the cheapest way to find out whether this is worth your evening.

Tested on: not independently tested. Every speed, VRAM, RAM, disk and predictability figure is author-reported in the project's own README, measured on the author's hardware, with no independent replication at the time of writing. The negative results on speculative decoding are also the author's own, and are reported here because publishing them is to the project's credit. Repository statistics were read from the GitHub API on the date below, when colibri stood at 38,125 stars.
Date checked: 2026-09-28

Prev Article
PageIndex
Next Article
DeepSeek Harness, 39 days later

Related to this topic: