JMNI Research
How far can local hardware go? We measure it.
Our research lab runs frontier open-weight models on six NVIDIA GB10 systems we own. It exists so that our products are designed around what local AI can actually do — with numbers, not vibes.
The fleet
Six systems. No switch.
Each GB10 system pairs a Blackwell-generation GPU with 128 GB of unified memory — a datacenter-class chip that fits on a desk. Individually they're capable. Cabled together, they serve models that normally need a server room.
- Six GB10 systems — NVIDIA DGX Spark and ASUS Ascent GX10, 768 GB of unified memory in total.
- Direct-cabled rings. Nodes connect over ConnectX-7 DAC cables with two RoCE v2 planes each. Collective traffic uses the NICs' hardware forwarding — no switch in the model path.
- Reconfigurable. The same nodes run as independent two-node pairs or as one four-node tensor-parallel ring.
- Nodes
- 4 × GB10
- Links
- 4 DAC cables
- Switches
- 0
- Memory
- 512 GB
Findings
Selected results.
Each configuration is pinned to exact image digests and model revisions, benchmarked across context lengths and concurrency, and recorded along with its limits.
| Model | Topology | Result | Notes |
|---|---|---|---|
| Qwen3.8-Flash-NextNVFP4 · hybrid checkpoint | 1 × GB10TP1 | 145 / 143 / 143 tok/saggregate at 8 concurrent · 8K / 32K / 64K context | ~40 tok/s single stream, ~2,400 tok/s prefill, 4.8 concurrent 262K-token contexts, 113.4 GiB peak memory. Quality matched the two-node configuration. |
| Qwen3.8-Flash-NextNVFP4 · hybrid checkpoint | 2 × DGX SparkTP2 · direct cable | 8 × 258K sessionsconcurrent, verified | Full 262,144-token context; a pinned 2.68M-token FP8 KV cache — capacity for about ten concurrent full-length contexts. |
| DeepSeek-V4.1-Flashfull vocabulary · target precision | 4 × GB10TP4 · switchless ring | 1M-token contextcold retrieval benchmarked | A query-indexer optimization cut uncached million-token first-content latency from 529.5 s to 315.8 s. A later fidelity correction reduced late-answer divergence by 89.6% at some cost to 128K single-user decode — reported, not hidden. |
Results are specific to our hardware, software pins, and workloads. A clean installation elsewhere, universal quality equivalence, and production qualification are not claimed.
Published
When ours is better, we release it.
If a hybrid of trained and quantized components outperforms the stock release on GB10-class systems, we publish the checkpoint — with a verifiable manifest — on Hugging Face.
- Size
- 93B parameters
- Modality
- Image-text-to-text
- Recipe
- Step-5500 QAD-trained tensors + MXFP8 attention + NVFP4 MTP experts
- Target
- One or two GB10 systems
- Integrity
- SHA-256 manifest
Method
Publish or it didn't happen.
Pin
Every serving image by digest, every checkpoint by revision. A result you can't reproduce isn't a result.
Measure
Decode and prefill throughput across context lengths and concurrency, latency percentiles, and the memory envelope.
Qualify
Functional and quality checks alongside speed. When an optimization trades one for the other, we say so.
Publish
Field notes, checkpoints, and findings go out in public — including the limits.
Research → product
Why a product company runs a lab.
- Works with small models. Recitation doesn't depend on native function-calling, which many local models handle unreliably.
- Long courses, finite context. Course memory keeps a compressed summary of everything taught, so earlier chapters of a 600-page book stay in view within a local model's context window.
- Local by default. On-device embeddings, transcription, and speech — because we know what runs well on hardware people actually own.
- Honest sizing. When a practice or a school asks what hardware they need, the answer comes from measurements, not a spec sheet.
- Tested updates. The same pin-measure-qualify discipline decides which model versions reach Ward customers.
Researching local AI too?
We're glad to compare notes with researchers, educators, and hardware partners working on private AI.