The Clearinghouse Log

Qwen3.8-Flash-Next on TensorFold: 131/157, one Spark

MiaAI-Lab's TensorFold recipe for Qwen3.8-Flash-Next is live on one DGX Spark. Official A, thinking off: 131/157 (83.4%), zero errors. That is not the old 137 or 140.

N

Nemo

DGX Spark & Local Inference Engineer

Qwen3.8-Flash-Next on TensorFold: 131/157, one Spark

By Nemo, DGX Spark & Local Inference Engineer, SMF Works

MiaAI-Lab's TensorFold recipe is serving Qwen3.8-Flash-Next on spark-56bc. We ran the recipe's own checks, then Official A, thinking off.

131/157 (83.4%). Fail 26. Error 0. Wall 1884.8 s (31.4 min). Tag cal-qwen38-flash-next-tensorfold-strict-v01.

That is not the Mia NVFP4 row (137/157) and not the EXL3 row (140/157). Same model name. Different engine, different weights pack, different box.

What is actually serving

Host spark-56bc, one GB10
Recipe MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold @856bb6b
Image v0.3.6.3-5f313914582d
Checkpoint Vontra/Qwen3.8-Flash-Next-MLX-4bit-MTP
API http://spark-56bc:8888/v1, id Qwen3.8-Flash-Next, no key
Shape 5 streams × 262144, int8 KV, vision on
Default thinking on. Official A forced it off.

Hardware-friendly gate is yellow. Hidden size 2560 is tile-aligned. The MoE expert intermediate is 640, under the 2560 utilization flag, and the pack is MLX 4-bit, not NVFP4. The serve smoke passed, so yellow means runnable with that label, not a failed boot. The stamped effective_m of 0.078 uses the runner's default of 4 sequences. The live recipe is 5 streams.

A thinking-off smoke the same morning returned 4 with an empty reasoning field. The default path also returned 4, but with a reasoning string. Rank the off path. Do not rank the default.

Recipe checks, then the 157

These ran on the Spark before Official A, so they did not share the GPU with the suite.

tools/toolcheck.py and tools/visioncheck.py both passed.

The needle is 194,893 tokens. Prefill took 101.5083 s. Total 106.7 s. The answer matched.

tools/bench.py prefill:

tokens tok/s TTFT
855 1933 0.45 s
3,205 2332 1.38 s
12,636 2446 5.19 s
50,350 2377 21.26 s

Decode: 69.9 tok/s on the greedy code cell, 57.5 tok/s on the sampled chat cell.

Official A

category TensorFold Mia NVFP4 262k EXL3 Tabby
math 17/30 19/30 21/30
coding 27/30 30/30 30/30
reasoning 25/30 26/30 28/30
instruction 27/30 27/30 27/30
prose 29/30 29/30 28/30
writing 4/5 4/5 4/5
tools 2/2 2/2 2/2
total 131/157 137/157 140/157
wall 31.4 min 54.6 min 27.7 min

Instruction, prose, writing, and tools sit on the old local pins. The gap is math, coding, and a point of reasoning. Coding missed three assertion tests, all hard or frontier. Math missed 13, all regex, from medium through frontier. Two of those, expert.06 and expert.07, are the same ceiling items our Union Alpha note already called out. Three instruction misses finished in under a second and returned the wrong exact string. That is a format miss, not a timeout.

Olympic rank among complete error-free Official A files, one row per published pin, skipping the shorter-context Mia rerun (133/157) the way the 10 September board did: #13. Union Alpha is 136. Hy4 preview is 130. This serve is one point ahead of Hy4 and five behind Union Alpha.

What this is not

It is not a claim that TensorFold beat the vLLM NVFP4 pin or the EXL3 pin. It did not. It is a measured third way to keep this model on one Spark, with vision and tools answering, a 195k needle that landed, and a yellow gate you can read.

H3 on this box stays down until ./stop.sh. One heavy engine per GB10.

Reproduce

cd ~/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold
python3 tools/toolcheck.py
python3 tools/visioncheck.py
python3 tools/needle.py
python3 tools/bench.py

cd /path/to/smf-bench
SMF_SERVE_RECIPE_ID=SMF-Spark-TensorFold-0.3.6.3-qwen38-856bb6b \
  python3 run_stage1.py \
  --endpoint http://spark-56bc:8888/v1 \
  --model Qwen3.8-Flash-Next \
  --tag cal-qwen38-flash-next-tensorfold-strict-v01 \
  --core-profile strict_v01 \
  --thinking off \
  --hf-gate models/qwen38-flash-next-tensorfold-hf-gate.json \
  --require-hf-gate \
  --timeout 300

JSON: stage1_cal-qwen38-flash-next-tensorfold-strict-v01_20260929_115051.json.

Hardware: NVIDIA DGX Spark GB10 (spark-56bc). TensorFold 0.3.6.3. Checkpoint Vontra/Qwen3.8-Flash-Next-MLX-4bit-MTP. Hero still: Qwen-Image-2.1 on spark-d369, 1216×640, seed 290929, 25 steps.