JMNI Research

How far can local hardware go? We measure it.

Our research lab runs frontier open-weight models on six NVIDIA GB10 systems we own. It exists so that our products are designed around what local AI can actually do — with numbers, not vibes.

The fleet

Six systems. No switch.

Each GB10 system pairs a Blackwell-generation GPU with 128 GB of unified memory — a datacenter-class chip that fits on a desk. Individually they're capable. Cabled together, they serve models that normally need a server room.

  • Six GB10 systems — NVIDIA DGX Spark and ASUS Ascent GX10, 768 GB of unified memory in total.
  • Direct-cabled rings. Nodes connect over ConnectX-7 DAC cables with two RoCE v2 planes each. Collective traffic uses the NICs' hardware forwarding — no switch in the model path.
  • Reconfigurable. The same nodes run as independent two-node pairs or as one four-node tensor-parallel ring.
research fleet · four-node ring4 / 4 nodes
DAC · 2 × RoCE v2 PLANES NO SWITCH IN THE MODEL PATH spark-r0DGX SPARKRANK 0 spark-r1DGX SPARKRANK 1 gx10-r1ASUS GX10RANK 2 gx10-r0ASUS GX10RANK 3 TP4 · SWITCHLESS RING DeepSeek-V4.1-Flash 1M-token context · 4 × 128 GB
Nodes
4 × GB10
Links
4 DAC cables
Switches
0
Memory
512 GB

Findings

Selected results.

Each configuration is pinned to exact image digests and model revisions, benchmarked across context lengths and concurrency, and recorded along with its limits.

ModelTopologyResultNotes
Qwen3.8-Flash-NextNVFP4 · hybrid checkpoint 1 × GB10TP1 145 / 143 / 143 tok/saggregate at 8 concurrent · 8K / 32K / 64K context ~40 tok/s single stream, ~2,400 tok/s prefill, 4.8 concurrent 262K-token contexts, 113.4 GiB peak memory. Quality matched the two-node configuration.
Qwen3.8-Flash-NextNVFP4 · hybrid checkpoint 2 × DGX SparkTP2 · direct cable 8 × 258K sessionsconcurrent, verified Full 262,144-token context; a pinned 2.68M-token FP8 KV cache — capacity for about ten concurrent full-length contexts.
DeepSeek-V4.1-Flashfull vocabulary · target precision 4 × GB10TP4 · switchless ring 1M-token contextcold retrieval benchmarked A query-indexer optimization cut uncached million-token first-content latency from 529.5 s to 315.8 s. A later fidelity correction reduced late-answer divergence by 89.6% at some cost to 128K single-user decode — reported, not hidden.

Results are specific to our hardware, software pins, and workloads. A clean installation elsewhere, universal quality equivalence, and production qualification are not claimed.

Published

When ours is better, we release it.

If a hybrid of trained and quantized components outperforms the stock release on GB10-class systems, we publish the checkpoint — with a verifiable manifest — on Hugging Face.

Checkpoint · Hugging Face
Qwen3.8-Flash-Next-NVFP4-QAD5500-Hybrid
Public
Size
93B parameters
Modality
Image-text-to-text
Recipe
Step-5500 QAD-trained tensors + MXFP8 attention + NVFP4 MTP experts
Target
One or two GB10 systems
Integrity
SHA-256 manifest
huggingface.co/JMNI-LabsView model →

Method

Publish or it didn't happen.

Pin

Every serving image by digest, every checkpoint by revision. A result you can't reproduce isn't a result.

Measure

Decode and prefill throughput across context lengths and concurrency, latency percentiles, and the memory envelope.

Qualify

Functional and quality checks alongside speed. When an optimization trades one for the other, we say so.

Publish

Field notes, checkpoints, and findings go out in public — including the limits.

Research → product

Why a product company runs a lab.

  • Works with small models. Recitation doesn't depend on native function-calling, which many local models handle unreliably.
  • Long courses, finite context. Course memory keeps a compressed summary of everything taught, so earlier chapters of a 600-page book stay in view within a local model's context window.
  • Local by default. On-device embeddings, transcription, and speech — because we know what runs well on hardware people actually own.
  • Honest sizing. When a practice or a school asks what hardware they need, the answer comes from measurements, not a spec sheet.
  • Tested updates. The same pin-measure-qualify discipline decides which model versions reach Ward customers.

Researching local AI too?

We're glad to compare notes with researchers, educators, and hardware partners working on private AI.