A storage hierarchy you can inspect
Memory becomes part of the runtime contract.
Locate the fused slice, lease it from the cache, read on miss, execute with AVX-512 VNNI, then release. The boundary is explicit at every step.
Sparse expert inference for the memory boundary
SparseFlow treats RAM, SSD and native CPU execution as one explicit runtime. Dense state stays close. Routed experts travel only when the route asks for them.
The useful idea is simple
Qwen3.6 carries a large expert bank, but each token activates a sparse route. SparseFlow makes that imbalance operational: resident dense state, bounded expert cache, and a native kernel at the point of use.
A storage hierarchy you can inspect
Locate the fused slice, lease it from the cache, read on miss, execute with AVX-512 VNNI, then release. The boundary is explicit at every step.
Bounded by design
256 MiB to 4 GiB cache options let Doctor match the runtime to currently available RAM.
See the admission modelMeasured, not imagined
Inspect token forwards, active experts, logical requests and physical reads without hiding the mechanism behind a screenshot.
Open the evidence viewThree ways to see the system
Hover or select a path. Each schematic describes a real runtime boundary; measured UI captures remain available in the repository README.
Measured, then explained
These are warm single-request W8A8 hybrid measurements on an Intel Xeon Gold 6248R host with 10 CPU threads and NVMe storage. They show the trade: resident execution avoids reads, while streaming keeps the working set bounded.
The chart is the formal Stage 7.5.6 matrix. The three performance cards below are Stage 7.6 follow-up measurements with a different runtime protocol; they are reported separately, not as one combined run.
Stage 7.6 follow-up · Resident / full working set
Stage 7.6 follow-up · Streaming / 4 GiB cache
Stage 7.6 follow-up · Streaming / 8 GiB cache
Quality boundary / 60 questions
59 / 60Resident vs streaming: 60 / 60 exact agreement. Native resident and native 4 GiB streaming matched each other on every prediction and choice total; character-normalized agreement with BF16 was 60 / 60.
Model-cold / 4 GiB cache
1.0227 tok/sThree independent cold processes measured 57.10 s TTFT median and no more than 9.75 GiB peak RSS. Cold and warm results are reported separately.
Full methodology: Stage 7.6 critical-path report, Stage 7.5.6 formal matrix, and the ten-day build log.
Start with the boundary
Read the limits, run Doctor, prepare the container, and watch the route move through your own machine.