The lab runs six GB10-class systems: NVIDIA DGX Spark and ASUS Ascent GX10 machines. Each one is a Blackwell-generation GPU with 128 GB of unified memory — a datacenter chip that fits on a desk. That's the whole point of this hardware class: small enough to own, big enough to be interesting.
Interesting, though, is not the same as easy. A pair of these systems can serve genuinely frontier open-weight models — if you can get the pieces to cooperate. That "if" is where most people stall, and it's where we spend our research time.
The ring
The configuration we care most about: instead of putting a switch in the model-traffic path, the nodes are cabled directly to each other over their ConnectX-7 links in a ring, and collective traffic is forwarded through the NICs' own hardware. In plain terms — every node can talk to every other at high speed, with no network box in the middle. Fewer parts, less to configure, less to fail.
The GB10 community has done remarkable work making these stacks real. We build on that work, measure it carefully, and write down what we find.
What runs here
Among others: GLM-5.3-Flash in NVFP4 across a Spark pair; Qwen3.8-Flash-Next on single systems and pairs, including a hybrid checkpoint we published ourselves; and DeepSeek-V4.1-Flash across a four-node switchless ring with million-token context. Selected results are on our research page.
Publish or it didn't happen
A performance claim without evidence is marketing. So every configuration we report is pinned to exact versions, measured across context lengths and concurrency, and written up with its limits — including when an optimization makes one thing faster and another thing slower.
That discipline carries straight into our products. When we say our products work well on local hardware, it's because we've measured what local hardware can do.