Sparse expert inference for the memory boundary

Move only the experts a token needs.

SparseFlow treats RAM, SSD and native CPU execution as one explicit runtime. Dense state stays close. Routed experts travel only when the route asks for them.

Qwen3.6 checkpoint anatomy showing 60 GiB routed experts and a 6.97 GiB dense resident core
A visible route is a better performance story than a hidden tensor copy.
QWEN3.6TOP-K ROUTINGINT8 EXPERTSAVX-512 VNNISSD STREAMINGROUTE TRACE QWEN3.6TOP-K ROUTINGINT8 EXPERTSAVX-512 VNNISSD STREAMINGROUTE TRACE

The useful idea is simple

Large models do not need one memory tier.

Qwen3.6 carries a large expert bank, but each token activates a sparse route. SparseFlow makes that imbalance operational: resident dense state, bounded expert cache, and a native kernel at the point of use.

SparseFlow request data flow from host runtime through cache and SSD to the native kernel

A storage hierarchy you can inspect

Memory becomes part of the runtime contract.

Locate the fused slice, lease it from the cache, read on miss, execute with AVX-512 VNNI, then release. The boundary is explicit at every step.

Bounded by design

Cache the hot route. Keep the rest on disk.

256 MiB to 4 GiB cache options let Doctor match the runtime to currently available RAM.

See the admission model
SparseFlow route trace showing selected experts and physical reads

Measured, not imagined

Every route leaves evidence.

Inspect token forwards, active experts, logical requests and physical reads without hiding the mechanism behind a screenshot.

Open the evidence view

Three ways to see the system

Pick a memory story.

Hover or select a path. Each schematic describes a real runtime boundary; measured UI captures remain available in the repository README.

SparseFlow resident memory path
Resident path / full working set

Measured, then explained

Less memory changes the curve, not the contract.

These are warm single-request W8A8 hybrid measurements on an Intel Xeon Gold 6248R host with 10 CPU threads and NVMe storage. They show the trade: resident execution avoids reads, while streaming keeps the working set bounded.

QWEN3.6-35B-A3BW8A8 NATIVE HYBRID32 GREEDY TOKENSFORMAL MATRIX / STAGE 7.5.6FOLLOW-UP CARDS / STAGE 7.6

The chart is the formal Stage 7.5.6 matrix. The three performance cards below are Stage 7.6 follow-up measurements with a different runtime protocol; they are reported separately, not as one combined run.

Formal decode throughput comparison across resident, streaming and generic offload paths
Formal Stage 7.5.6 matrix. The highlighted warm 4 GiB streaming cell is 1.41 tok/s at 9.87 GiB RSS.
Correctness gate showing exact resident and streaming agreement across 60 questions
Resident vs streaming: 60 / 60 exact agreement.

Stage 7.6 follow-up · Streaming / 4 GiB cache

Low-memory path

1.4128 tok/s
Peak RSS
9.87 GiB
Expert reads
20.76 GiB
TTFT
12.36 s

Stage 7.6 follow-up · Streaming / 8 GiB cache

More locality

1.4944 tok/s
Peak RSS
13.94 GiB
Expert reads
13.84 GiB
TTFT
12.46 s

Quality boundary / 60 questions

59 / 60

Native vs BF16

Resident vs streaming: 60 / 60 exact agreement. Native resident and native 4 GiB streaming matched each other on every prediction and choice total; character-normalized agreement with BF16 was 60 / 60.

Model-cold / 4 GiB cache

1.0227 tok/s

Cold storage is a separate cost.

Three independent cold processes measured 57.10 s TTFT median and no more than 9.75 GiB peak RSS. Cold and warm results are reported separately.

Full methodology: Stage 7.6 critical-path report, Stage 7.5.6 formal matrix, and the ten-day build log.

Start with the boundary

Run the model where your memory actually is.

Read the limits, run Doctor, prepare the container, and watch the route move through your own machine.