Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

TensorFold - Inference

GLM-5.3-Flash on Two DGX Sparks: The TensorFold Recipe

Two DGX Sparks, a cable between them, and GLM-5.3-Flash answering on an OpenAI-compatible endpoint with its full million-token window open. The repository publishes enough numbers to decide before you buy the hardware, and one licence detail that decides it for commercial users.

License Apache 2.0
License Apache 2.0
TL;DR
  • Two DGX Sparks linked by ConnectX-7 serve the full 1,048,576 token window on one endpoint
  • 108.8 tok/s prose and 227.9 structured across four concurrent requests, community-reported
  • Every published figure used a drafter licensed non-commercially; the commercial path is unmeasured

Two DGX Sparks, a cable between them, and a trillion-token-class model answering on an OpenAI-compatible endpoint with its full million-token window open. That is what this recipe builds, and it publishes enough numbers that you can decide whether it is worth the hardware before you buy any.

It is also the first TensorFold setup we have seen written up as a complete production build rather than a benchmark. We covered TensorFold and its two rivals when the interesting part was the exactness claim. This is what the project looks like when someone actually ships with it.

One thing to know before you start pricing Sparks: the fast path, as configured, is not licensed for commercial use. More on that below, because it is the detail most likely to bite.

The hardware

Two NVIDIA DGX Sparks, the GB10 variant, 128 GB of unified memory each. They are linked directly by their ConnectX-7 QSFP ports with a cable, and talk over RoCE v2. No switch.

If you have not met a Spark, our hardware comparison covers where it sits against an M4 Max Mac Studio and a Ryzen AI Max 395.

The floors you need to clear on each machine:

ResourceRequirement
Free memory at startupAbout 110 GB per Spark
DiskAbout 205 GB: ~176 GB checkpoint, ~2.3 GB drafter, ~25 GB image
Preflight checks180 GB free for the download, 35 GB under Docker's root
SoftwareDocker with the NVIDIA container runtime, rsync, passwordless ssh to the worker

You can avoid the second copy of the weights. Set WORKER_WEIGHTS=nfs and the worker reads them from the head instead of keeping its own 176 GB.

What is actually running

The model is GLM-5.3-Flash in EXL3, routed experts quantized to 4 bits a weight and BF16 elsewhere, about 176 GB on disk. The dense weights are configurable as 4-bit, FP8 or BF16.

The engine is TensorFold v0.6.0 plus 53 patches, one rank on each Spark, inside NVIDIA's PyTorch container. The patches are where most of the work is:

  • DFlash2 and copy drafts
  • 4-bit dense weights and an FP8 KV cache
  • Faster prompt kernels
  • A one-shot RoCE all-gather between the two Sparks
  • Several requests sharing one cache pool
  • Vision, tool calling, /tokenize and /metrics

Each request can use the full 1,048,576 token context. Four concurrent requests share an FP8 KV pool of about 2.9 million tokens.

The numbers

All community-reported from the repository. Measured on two Sparks at the default configuration, four streams, FP8 KV cache, 4-bit dense weights, DFlash2 plus copy drafts, vision on, with GPU clocks capped at 2,200 MHz, using sparkDash over the OpenAI API from another machine on the network.

Decode, aggregate across concurrent requests and per request:

ConcurrentProsePer requestTTFTStructuredPer requestTTFT
160.4 tok/s60.4170 ms114.7 tok/s114.7149 ms
279.2 tok/s40.4269 ms147.6 tok/s77.1295 ms
389.5 tok/s30.6330 ms196.3 tok/s67.4317 ms
4108.8 tok/s27.9340 ms227.9 tok/s61.1415 ms

Structured output runs roughly twice the prose rate throughout, which is worth noting if your workload is tool calls and JSON rather than paragraphs.

Prefill holds close to 2,000 tok/s up to about 65k tokens, then falls away as the window grows:

PromptPrefillTime to first token
8,219 tokens1,952.2 tok/s4.21 s
32,790 tokens1,978.9 tok/s16.57 s
65,563 tokens1,942.5 tok/s33.75 s
131,099 tokens1,837.9 tok/s71.33 s
262,170 tokens1,641.7 tok/s159.69 s
981,841 tokens1,015 tok/s967 s, needle found

That last row is the honest shape of a million-token window. Sixteen minutes to first token. The capability is real and so is the wait.

Which makes the prompt reuse numbers the most practical thing in the table:

PromptFirst timeNext time
An identical 64k prompt, sent again34 sunder 0.07 s
A new conversation, same 7.9k system prompt4.24 s0.13 s

On quality, measured on a build with identical replies: GSM8K 98.0% over 250 problems, HumanEval 97.6% over 164.

And in keeping with the exactness theme that made TensorFold interesting in the first place, the repo checks that concurrency does not change answers. Replies served four at a time matched the same requests served one at a time in 11 of 11 staggered cases and 11 of 11 sent as a burst.

The licence stack

This is the part to read twice.

Five licences apply to one running server, and they do not all point the same way:

ComponentLicence
The recipe repositoryApache 2.0
TensorFold v0.6.0Apache 2.0, with MIT notices for earlier code
GLM-5.3-Flash base weightsPer its model card, MIT when Z.ai released it
The EXL3 checkpointShapleyMcg License 1.0, attribution required
The DFlash2 drafterCC BY-NC-ND 4.0, non-commercial only
The container baseNVIDIA Software License Agreement

Two of those deserve attention.

The drafter is non-commercial. DFlash2 is CC BY-NC-ND 4.0. The repo is upfront about it and ships an alternative: set DRAFTER=mtp and it serves using the checkpoint's own MTP head instead, with no DFlash2 involved. Commercial licensing is available from the authors.

Here is the catch worth stating plainly: every performance figure above was measured with DFlash2 in the loop. The repository does not publish a second set of numbers for the mtp path. If you are deploying commercially, the configuration you can legally run is not the one that was benchmarked, and you should expect to measure it yourself.

The checkpoint licence has an unusual clause. The EXL3 weights are under the ShapleyMcg License 1.0, which requires attribution and states that it "grants no rights to the person known as '0xSero.'" A licence that excludes a named individual is rare enough to be worth knowing about before you build a product on it.

Gotchas

Vision costs window. With vision enabled the tower takes about 1.8 GB on rank 0, which costs roughly 66k tokens of context in dense FP8 mode.

First run is long. The initial launch pulls or builds the image, downloads 176 GB of checkpoint plus the drafter, copies to the worker and compiles CUDA kernels once per image. Later starts take 2 to 6 minutes to load about 80 GiB of weights on each Spark.

Concurrency has open issues. Issues #12 and #13 track prompt eviction with TF_GLM_MULTI_LONE=1. Four open issues at the time of writing against 122 stars and 10 forks, so this is a young project with an attentive author rather than a settled one.

Standing it up

The sequence is short, which is the point of the repository:

git clone https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold.git
cd GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
cp scripts/local.sh.example scripts/local.sh
# edit scripts/local.sh and set WORKER=user@<worker-ip>
./start.sh

start.sh prepares both ranks, runs a smoke test through them and prints the endpoint. Then it is an ordinary OpenAI-compatible server:

curl http://<head>:8888/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"GLM-5.3-Flash-EXL3","messages":[{"role":"user","content":"hello"}],"max_tokens":2000}'

./stop.sh brings both ranks down.

Who this is for

Two DGX Sparks is a serious outlay, and this recipe does not make it cheaper. What it does is remove the part that usually stops people: making two of them behave like one inference target, with a context window nobody else is offering at this scale locally.

It fits if you need a million-token window on hardware you own, if your data cannot leave the building, or if you are running structured workloads where 228 tok/s across four streams is the shape of the job.

It does not fit if you want one model on one box. A single Spark will not hold this checkpoint, and the recipe is explicitly built for the pair.

Sources and further reading

Ten minutes, and no hardware required: open the repository's README and read the Performance section against your own workload. If your prompts are long and repeated, the reuse numbers matter more than the decode rate. If they are short and varied, the opposite. That comparison tells you whether two Sparks would change anything for you, and it costs nothing to make.

Tested on: not independently reproduced. No two DGX Sparks here. Every figure above is community-reported from the project's README, measured by its author with sparkDash at 2,200 MHz clocks with DFlash2 enabled, and is reprinted rather than verified. The licence terms were read from the repository's own LICENSE and NOTICE files on the date below.
Date checked: 2026-10-01

Prev Article
n8n MCP Features Explained: Five Things, One Confusing Name
Next Article
How to create Logos with Midjourney

Related to this topic: