DeepSeek Harness, 39 days later

In August we audited DeepSeek Harness and listed five things to watch. Five weeks and 66,000 stars later we re-ran the same queries. One prediction landed, and only halfway.

Colibri

Colibri streams MoE experts from NVMe so frontier models never have to fit in RAM. Kimi K3 on 32 GB, GLM-5.2 on 16 GB, pure C, Apache 2.0. And about 1.8 tokens a second.

PageIndex

PageIndex replaces vector search with a tree index an LLM navigates. How it really works, why the headline benchmark is 90.7% strict, and how to run it against a local model.