Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

Strata - Inference

Strata: Run a 125B Model on a 12 GB Gaming GPU

Strata runs Qwen3.8-Flash-Next, a 125B MoE model, on a 12 GB gaming GPU at 94 tokens per second. How the three-tier memory split works, what the self-reported numbers do and do not show, and the licence detail that decides whether you can ship it.

License MIT
License MIT
TL;DR
  • 125B total, but only 6B activates per token, which is why 12 GB of VRAM is enough
  • Three tiers: hot experts in VRAM, all experts in RAM, the 51B embedding table on SSD
  • 94 tok/s on an RTX 5070, self-reported, with no quality benchmark published alongside
System Requirements
RAM32 GB minimum, 64 GB runs every quantization
GPUNVIDIA RTX 20/30/40/50 or AMD RX 7000/9000 series; RTX 5070 about 94 tok/s
VRAM12 GB minimum; a 24 GB RTX 3090 is estimated at 100 to 140 tok/s

A 125-billion-parameter model on a 12 GB gaming card. Not a distilled version of it, not a 7B wearing its name: the actual weights, answering at 94 tokens per second.

That is what Strata claims, and the claim is ten days old. The repo went up on 24 September 2026 and has 10,573 stars and 930 forks at the time of writing. It is MIT, it installs with one click on Windows or Linux, and it serves an OpenAI-compatible API on localhost.

Here is how it works, what the numbers actually say, and the licence detail that decides whether you can use any of it at work.

Why a 125B model fits at all

The trick is not compression. It is that this model was never dense to begin with.

Strata runs Qwen3.8-Flash-Next, a mixture-of-experts model. A dense model runs every parameter for every token. An MoE model has many small specialist networks and a router that wakes only a few per token.

The numbers from the model's own config.json:

PropertyValue
Total parameters125B, plus 51B of n-gram embeddings and 4B of MTP
Activated per token6B
Experts per layer512
Layers48
Experts woken per token10 routed, plus 1 shared
Context262,144 native, extensible to 1,000,000
InputsText, images and video

You will see "24,576 experts" in the coverage. That is 512 per layer times 48 layers, counting every expert instance in the model. The model card's 512 counts them per layer. Both are correct, they are just different units, and it is worth knowing which one someone means.

The part that matters for your GPU: 6B of 125B does the work for any given token. You do not need all the weights in VRAM at once. You need the right ones, at the right moment.

Three tiers, and the SSD is doing real work

Strata splits the model across the memory you already have, by how often each piece gets touched.

TierWhat lives there
GPU VRAMThe few thousand experts the router asks for most
System RAMEvery expert. If the GPU does not have the one needed, the CPU handles that token
SSDThe 51B n-gram embedding table, read in small rows per word

That third row is the clever one. Those 51B parameters are a lookup table, not a matrix multiply. A lookup only needs the handful of rows for the words actually in front of it, and an SSD is perfectly good at fetching small rows. So a third of the model's bulk never competes for RAM at all.

On top of that, Strata uses the model's own multi-token prediction head for speculative decoding: a small fast path guesses the next few tokens, the full model checks them in one pass, and accepted guesses come free. The project reports that as a 1.6x to 1.8x multiplier.

If this shape sounds familiar, it should. We covered Colibri streaming MoE experts from disk a week earlier. Same insight, different tier: Colibri reaches down to storage for the experts themselves, Strata keeps experts in RAM and sends only the embedding table to disk.

The numbers

These are the project's own measurements, not ours. Strata publishes them with the engine version attached, which is more disclosure than most projects manage.

NVIDIA RTX 5070 (12 GB), Ryzen 5 7600, 64 GB DDR5-5200, Windows. Output measured on a 4K prompt, prompt processing on a 32K one.

QuantizationOutput tokens/sPrompt tokens/s
Q2_0942,650
IQ2_XS792,090
IQ3_XXS621,750
IQ3_S531,620
Coder (IQ1_M)552,180

AMD RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM.

QuantizationOutput tokens/sPrompt tokens/s
Q2_0601,160
IQ2_XS521,110
Coder (IQ1_M)441,420

Two caveats the project states itself, and you should carry them with the numbers. The Q2_0 row was measured on engine 0.1.36 while the others are 0.1.26, so the table mixes versions. And output speed moves by several percent run to run, because speculative decoding goes faster when more of its guesses are accepted, which depends on what the model is writing.

Speed also falls as context grows. On the same RTX 5070 at Q2_0, the longer tables show 93 tokens/s at 4K dropping to 73.7 at 128K and 60.3 at the full 262K window.

An RTX 3090 with 24 GB is estimated at 100 to 140 tokens/s, but that one is an estimate rather than a measurement.

Installing it

You need a 12 GB or larger NVIDIA RTX 20/30/40/50 card or a supported AMD Radeon, 32 GB of RAM, and about 80 GB of free disk. 64 GB of RAM runs every quantization; 32 GB restricts you to the smaller ones.

On Linux:

git clone https://github.com/Niko1221/Strata.git
cd Strata
./setup.sh

On Windows, unzip the project and run START-HERE.bat. Either way it asks which quantization you want, how much context, and whether you need image input, then fetches the weights.

What you get is a local server on port 8080 speaking two API dialects:

# OpenAI-compatible
http://127.0.0.1:8080/v1

# Anthropic-compatible
http://127.0.0.1:8080/v1/messages

That second one matters more than it looks. An Anthropic-compatible endpoint means coding agents that expect that API can be pointed at your own GPU by changing a base URL.

To expose it on your network:

./setup.sh --host 0.0.0.0 --api-key your-secret-here

The licence stack

This is the part to read before you put it in a product, because the permissive licence everyone quotes is not the one that governs the weights.

ComponentLicence
Strata, the engineMIT
llama.cpp / ggml underneathMIT
ISTA-DASLab GSQ-RCO GGUF quantizationsApache 2.0
Qwen3.8-Flash-Next base weightsqwen-community-1.0

The engine is MIT and the quantization repo is Apache 2.0, so a quick glance suggests you are clear. You are not, quite. A quantization is a derivative of the base weights, and the base weights ship under Qwen's own community licence rather than Apache or MIT. Read that licence against your use case before shipping anything on top of it.

Worth noting who did the compression work: the GSQ-RCO quantizations come from ISTA-DASLab, whose method we covered when it launched, with other sizes from UkisAI and Unsloth. On Hugging Face the ISTA-DASLab GGUF repo is pulling about 1.89 million downloads a month, more than the base model's 1.48 million. People are not downloading these weights to fine-tune them. They are downloading them to run.

Gotchas

  • One request at a time. This is a single-user engine. It is not a serving stack and does not pretend to be.
  • The first start is rough. Expect the machine to stall for one to three minutes while 35 to 55 GB loads into RAM.
  • Cold prompts are slow. Roughly a minute per 30,000 tokens on a first read, much faster afterwards.
  • The big quantizations need the RAM. IQ3_XXS and IQ3_S carry 43 and 50 GB of experts, so 262K context on a 64 GB machine runs out of memory. Setup caps you at 128K with those.
  • 4-bit KV cache has a real cost. It is optional and about 4% faster at 128K, but the project measures perplexity rising 8 to 12% on long documents. 8-bit stays the default for good reason.
  • AMD is newer than NVIDIA. The HIP backend works and is benchmarked, but it arrived later and image input is not available on it yet.
  • Q2_0 is a 2-bit quantization. The fastest row in that table is also the most compressed. Quality at 2 bits is not the same as quality at 4, and the project publishes no quality benchmark alongside the speed ones.

That last point is the real gap. Every number here measures how fast tokens come out, and none measures how good they are. For a quantization this aggressive, that is the question a reader should want answered, and it is not in the repo.

Who should run this

If you have a 12 GB card and 64 GB of RAM sitting in a desktop, this is the cheapest path to a frontier-class open model on your own hardware that currently exists. Nothing leaves the machine, there is no per-token bill, and the Anthropic-compatible endpoint means your existing agent tooling can use it.

If you are serving more than one person, look elsewhere. Single request at a time is a hard architectural limit, not a setting.

And if you are evaluating this for work, start with the licence table above rather than the speed tables.

The honest summary: a ten-day-old project, moving fast, documenting itself unusually well, with 10,573 stars and 260 open issues to match that pace. Treat the version numbers as load-bearing.

Sources and further reading

Ten minutes: check your RAM and your VRAM, then pick your row in the first table. 12 GB and 32 GB of RAM puts you on IQ2_XS, 64 GB opens up every size. If both numbers clear the bar, the installer does the rest while you make coffee.

Not independently tested. Every figure above is reported by the Strata project on the hardware it names, and we have not reproduced it. The model, architecture, parameter counts and licences were verified directly against the Hugging Face model card, its config.json and the GitHub repository on the date below.
Date checked: 2026-10-04

Prev Article
Omarchy: The Linux Desktop Built for AI Agents
Next Article
Pinokio Computer

Related to this topic: