SparseFlow · Build Log · 2026
Back to SparseFlow Pages

July 13 → July 23 · a 16 GB laptop → the same 16 GB laptop

Running a 67 GB MoE on a 16 GB Laptop

A ten-day build log: what we tried, what the numbers said, where our assumptions broke, and how we actually got Qwen3.6-35B-A3B running on the same 16 GB Lenovo we started on.

SparseFlow Public Alpha · code commit 36aa619 · experiment host: Intel Xeon Gold 6248R, 125 GiB RAM, 2×V100S (unused for inference), AVX-512 VNNI · laptop: 15.85 GiB RAM, 6-core CPU, no GPU

Where this starts

On July 13, I opened a Codex session on my laptop. It’s a pretty ordinary machine: 16 GB of RAM, a 6-core CPU, and no GPU that could usefully run this model. I started by feeding it a GitHub link to JustVugg/colibri—a project that somehow streams a massive 744B MoE directly from disk. I wanted to understand how they pulled that off, and more importantly, whether the same idea could run a smaller MoE on my own hardware.

We picked Qwen/Qwen3.6-35B-A3B as our target. It is a large model: 35B total parameters, about 3B active per token, 40 layers, 256 routed experts per layer, top-8 routing, one shared expert, Gated DeltaNet, Gated Attention, plus vision and MTP components. The full BF16 checkpoint is 67 GiB. A direct from_pretrained on a 16 GB laptop would run out of memory before the load completed.

My first sensible decision was actually not to develop on the laptop. I rented a high-memory CPU box for the coding and heavy experiments, keeping my laptop as the "reviewer's seat." At the time, I did not expect that ten days later, this same laptop would become the most convincing proof point.

To build this efficiently, I set up a "two-seat" workflow. I used a rented Xeon box (125 GiB RAM, fast NVMe) as my Implementation Seat to write the runtime and run heavy experiments. Meanwhile, my local laptop acted as the Review Board: I pulled clean commits here, parsed raw JSON, and carefully checked performance claims before merging them. This separation made the constraint visible early: if a feature worked on the big Xeon box but choked on the laptop's limited hardware, it did not pass.

1. The premise

If you inspect the Safetensors header—a small metadata block whose length is stored in the first 8 bytes—you can map the model without loading the payload. Once the tensors are grouped by responsibility, the memory problem becomes surprisingly concrete:

Qwen3.6-35B-A3B checkpoint anatomy The 67 GiB checkpoint splits into 60 GiB of routed experts, loaded on demand, and about 7 GiB of dense weights, kept resident. Vision weights are excluded. Qwen3.6-35B-A3B · checkpoint anatomy 66.97 GiB total · 1045 tensors · split by responsibility, not by tensor type LOADED ON DEMAND routed experts 40 layers × 256 experts · fused gate_up + down_proj 60.00 GiB ALWAYS RESIDENT 2.44 1.89 1.52 dense resident core attention · embed · lm_head · norms 6.97 GiB + 0.83 GiB vision (excluded from text-only Public Alpha) 60 GiB on demand · 7 GiB always
Figure 1. The 60 GiB of routed experts is our target for disk streaming. The ~7 GiB of dense weights is what we absolutely must keep in RAM.

Three numbers from this map became the core constraints of the project:

Once I wrote down the ratio—60 GiB on demand, 7 GiB always—the direction was clear. The hard part was turning that idea into a runtime: a correct locator, a bounded cache, leases that protect active tensors, an INT8 container, and a native AVX-512 VNNI kernel.

2. The architecture

We settled on a three-layer architecture with one golden rule: the core engine must be completely decoupled from Qwen. The host runtime handles attention, Gated DeltaNet, KV states, tokenization, and generation loops. SparseFlow's job is strictly focused: expert lookup, hot cache management, streaming I/O, and running our native W8A8 kernel. They talk through a very clean, narrow interface.

SparseFlow request data flow The host runtime asks SparseFlow for routed experts. The cache returns a pinned view on a hit, or the request flows through the locator and provider to SSD on a miss before the native W8A8 kernel runs. The kernel sends a routed output to the next-layer input, which returns to the host runtime through a dedicated right-side corridor. HOST RUNTIME · Transformers + PyTorch attention · Gated DeltaNet · KV cache · tokenizer · sampling · generation loop ① router picks top-8 routed experts (per layer) ② calls SparseFlow for those expert tensors ② SPARSEFLOW CORE · model-independent ③ ExpertCache (per-layer LRU · byte budget · leases) hit → return cached tensor view + pin lease HIT cached view + lease pinned MISS ④ ExpertLocator · fused-tensor byte range + dtype/shape, validated derives slice offset/length from safetensors header (no payload read) ⑤ Provider · pread INT8 payload from container resident = all experts preloaded at init · streaming = read on miss pread on miss SSD / NVMe INT8 · 30.078 GiB canonical-int8-v1 ⑥ Native W8A8 kernel · AVX-512 VNNI dpbusd · row-sum ⑦ routed output ROUTED OUTPUT next layer input RAM (resident) dense ~7 GiB + expert cache · 4 / 8 GiB budget — request — hit — miss — storage ┄┄ return
Figure 2. What happens when a decoder layer needs its experts. The host calls SparseFlow at ②; the cache either returns a pinned view on a hit, or falls through the locator → provider → SSD pipeline on a miss.

ExpertLocator: An expert is just a byte range

Qwen3.6 stores its routed experts in a fused layout. Expert 17 at Layer 0 isn't a standalone file; it's a specific slice along axis 0 of a giant fused tensor. The locator's job is to calculate the exact byte offset, length, dtype, and shape from the metadata header, and validate everything (expert counts, slice bounds, etc.) before we touch the disk. On the real model, Layer 0 Expert 17 resolves to exactly 6,291,456 bytes (6.00 MiB) across two shards, which we verified via SHA-256.

ExpertCache: Leases and byte budgets

The cache is a per-layer LRU with an optional global byte budget. The trickiest part here was implementing a lease system: when a forward pass pins an expert, the cache is forbidden from evicting it until the forward pass is done. If the cache is completely full of leased experts and we hit another cache miss, the provider falls back to returning a transient payload instead of corrupting an active tensor view.

The lease system wasn't in my original design. I had to add it after a brutal bug where a tensor view got silently evicted and overwritten mid-forward pass by a subsequent miss in the same layer. Now, we pin exactly what the current forward pass is using, release them in a finally block, and assert that the active lease count is exactly zero at the end of every token generation.

INT8 container: canonical-int8-v1

We packed everything into one .sfi file per layer: expert-major row-major S8, per-output-channel symmetric quantization, FP16 scales, 4 KiB alignment, and checksums everywhere. The conversion was processed one expert at a time, meaning we never had to materialize a giant fused layer in memory:

MetricValue
Routed BF16 source60.000 GiB
INT8 physical storage30.078 GiB (50.13%)
Experts / layers10,240 / 40
Conversion wall time373.03 s
Conversion peak RSS630.84 MiB
Max abs dequant error (samples)0.00201

We also precomputed the INT32 row sums into a sidecar file canonical-int8-exec-v1 (125.99 MiB), making runtime row-sum preparation practically instant.

3. Correctness first

Before optimizing for speed, we set a strict rule: the resident path (all in RAM) and the streaming path (from disk) must produce bit-identical outputs. No "close enough", no epsilon tolerances. We compared the full 248,320-way vocabulary logits at every single step (both prefill and decode) using SHA-256 hashes.

Correctness gate · resident and streaming paths must produce identical outputs Two storage paths feed the same eager expert kernel. A SHA-256 exact verifier compares logits, routes, IDs, and text at every step. The result is an exact match across all 60 quality questions with maximum logit delta zero. CORRECTNESS GATE SHA-256, every step, full 248,320-vocab RESIDENT PATH 60 GiB experts preloaded in RAM expert I/O = 0 RSS 66.4 GiB STREAMING PATH 4 GiB LRU cache experts read on miss from INT8 container RSS 9.6 GiB EXACT-EQUAL VERIFIER EXACT MATCH 60 / 60 MAX LOGIT DELTA 0
Figure 3. Two completely different storage paths feed the same eager expert kernel and produce bit-identical outputs.
Going for bit-exact outputs was a painful choice. The easy way out would have been "as long as the difference is under 1e-3, we're fine"—and for actual use, it is. But if we allowed even a tiny numerical drift, every benchmark afterward would be haunted by the question: "Is this actually working, or is it silently degrading?" When our first Layer-0 run returned max_abs_error = 0, I finally stopped worrying about correctness and started focusing on speed.

To prove this, we ran a 60-question evaluation set (HellaSwag, ARC-Challenge, MMLU, 20 questions each, seed=1234) comparing BF16 resident, W8A8 native resident, and W8A8 native streaming with a 4 GiB cache. The native resident and streaming paths matched perfectly on every single prediction, choice, and token log-likelihood. The streaming path read 767.36 GiB of expert payload over the entire run while strictly respecting the 4 GiB cache limit—and gave the exact same answers.

Formal matrix · Stage 7.5.6 · Xeon Gold 6248R · 10 threads

4. The numbers (on the big box)

Formal Stage 7.5.6 protocol: Xeon Gold 6248R, 10 CPU threads (we tested 1, 5, 10, and 20; 10 threads gave 2.934 tok/s, while 20 threads actually degraded to 2.837 due to overhead), one fixed prompt, 32 greedy tokens, one warmup run, and three measured runs. For the 4 GiB streaming path, we used three independent, cold-started processes to ensure no OS page-caching was cheating the results.

Decode throughput and RSS by path The chart compares eight paths. Bars stop before a dedicated value column, so each numeric result remains legible. The highlighted warm 4 GiB streaming path reaches 1.41 tok/s. Decode throughput by path experiment host · Xeon Gold 6248R · 10 threads · greedy · 32 tokens · warmup + measured runs TOKENS / SECOND · higher is better VALUE 0 1.0 2.0 3.0 Generic offload (warm) 0.028 Generic offload (cold) 0.003 BF16 streaming S1 8 GiB 1.08 W8A8 streaming 4 GiB (cold) 1.02 W8A8 stream · 4 GiB warm 1.41 W8A8 streaming 8 GiB (warm) 1.49 W8A8 native resident 2.49 BF16 resident (ceiling) 3.43 Bars end before the value column · highlighted path: 1.41 tok/s at 9.87 GiB RSS
Figure 4. Formal Stage 7.5.6 matrix. The streaming 4 GiB warm cell delivers 41% of the BF16 resident ceiling at 15% of the RSS. Versus a generic disk-offload baseline, it is 39.3× faster. Cold-cache streaming drops to 1.02 tok/s. We report both, never conflated.
PathDecodeTTFTRSSRead/token
BF16 resident (ceiling)3.43 tok/s4.67 s65.5 GiB0
W8A8 native resident2.49 tok/s8.25 s35.8 GiB0
W8A8 streaming 4 GiB (warm)1.41 tok/s12.36 s9.87 GiB355 MiB
W8A8 streaming 4 GiB (cold)1.02 tok/s57.1 s9.75 GiB355 MiB
Generic offload (warm)0.028 tok/s38.6 s5.8 GiBn/a

The 9.2 GiB moment

Stage 7.1, July 15. A fresh process generated four tokens from the 67 GiB model and peaked at 9.216 GiB RSS. The generated text was "Here's a thinking". The process never allocated the 60 GiB of expert tensors, not even temporarily.

I remember staring at that 9.216 GiB number for a while. Before Stage 7.1, "streaming experts" was a plausible design with good math behind it. After that run, it was a real runtime that fit in roughly a quarter of the model's size. The caveat mattered too: Linux page cache is outside process RSS, so we did not claim it as resident model memory. The process-level number was real and reproducible. The question had shifted from "can this work?" to "where does it break?"

5. Where it broke (and what we learned from stopping)

The negative results turned out to be some of the most useful results in the project. Each one was measured, gated, and kept out of the default path.

OptimizationHypothesisMeasuredVerdict
Pure fused decode kernel Faster expert micro-benchmark means faster end-to-end decode 0.886× legacy end-to-end NO-GO as default
Shared streaming batching Route union at B=8 should save shared-cache reads B=8 / 4 GiB union reads 1.205× more than round-robin NO-GO
lm_head native rewrite lm_head looks large enough to justify a rewrite About 5% of decode NO-GO (below 10% gate)
Deterministic current-layer I/O pipeline Overlap expert I/O with shared-expert compute 0.90 vs 1.21 tok/s synchronous NO-GO as default
DeltaNet projection fusion Small projection fusion should remove a visible decode cost Less than 1% of total decode NO-GO

The micro-benchmark trap

The pure fused decode kernel was the instructive failure. Its isolated eight-expert micro-benchmark was faster than the canonical per-expert path. Paired same-process AB/BA tests on the full runtime measured the fused path at 0.886× the legacy decode. It reduced routed-MoE time, but changed CPU frequency, thread-pool behavior, and memory access patterns enough to lose end-to-end throughput.

The micro-benchmark said "faster." The full-runtime AB/BA test said "not here." That is the part I want to remember: a kernel change perturbs everything around it. The fused kernel stayed available for diagnostics, while the canonical path kept its batch-one fast path.

When batching makes things worse

Stage 7.7 tested whether multiple concurrent requests could share a streaming cache and amortize reads. The route overlap was real: at B=8, the union of selected experts was 1.749× smaller than the sum of independent selections. The grouped kernel was exact and reached a 1.707× aggregate speedup at B=4. The failure was in the shared-cache gate:

Batch / CacheUnion hit rateRound-robin hit rateUnion / RR reads
B=4 / 4 GiB57.7%72.5%1.012×
B=8 / 4 GiB8.1%56.4%1.205×
B=4 / 8 GiB78.9%86.1%1.00×
B=8 / 8 GiB67.4%81.3%1.00×

B=8 / 4 GiB was the failure: 122.4 GiB union versus 101.6 GiB round-robin. At 8 GiB the result was close to neutral, too weak to ship a scheduler. The decision was explicit: resident-only fixed-cohort batching is implemented; shared streaming batching is disabled and labeled a known limitation.

This was the negative result I most wanted to keep. The route overlap existed, the grouped kernel worked, and the invariants passed. The goal still failed: reads did not go down. We kept the operator, disabled the scheduler, and labeled the limitation in the release instead of turning an interesting result into a product claim.

6. The laptop closes the loop

On July 23, ten days after the project started on the same machine, I closed most other applications on the 16 GB Lenovo, ran doctor with the real INT8 container, and got ready=true. Available RAM was 9.95 GiB, the minimum was 8.86 GiB, and the headroom was 1.08 GiB. The status was warn. Then came a two-token generation:

256 MiB expert cache
model load:       7.17 s
prefill:         21.83 s
single decode:    2.23 s   (~0.45 tok/s)
generated:       "Here's"
expert reads:   ~7.77 GiB
Two tokens: "Here's." The same two tokens the experiment host had generated ten days earlier at 9.2 GiB RSS, now on the laptop where the project started. A 67 GiB MoE had completed a real INT8 forward pass on a 15.85 GiB Windows machine with no GPU. It was slow and tight on memory, not something I would recommend for daily use. But it was real and reproducible on the kind of machine a reader of this post may own. I had been focused on the big-box numbers and nearly missed that this was the point.

The right framing is: experimental laptop preset, 0.45 tok/s, about 1 GiB headroom, and other applications closed. It is a proof of concept. The point is not that a 16 GB laptop is suddenly a comfortable 35B inference machine; the point is that the memory boundary is no longer an absolute wall.

7. Boundaries, and two bugs caught late

Here is what the Public Alpha is, and is not:

PathStatus
INT8 native resident hybridStable baseline
INT8 native single-request streaming S1 LRUStable low-memory
INT8 native resident grouped fixed cohortExperimental opt-in
Shared streaming batching / subcohortDisabled, known limitation
Pure fused decodeDiagnostics only

Generation is greedy-only. Conversation is stateless: the server replays the full message list rather than persisting KV state. AVX-512 VNNI is required. The host runtime still owns attention, Gated DeltaNet, KV state, tokenization, and the generation loop. SparseFlow is not yet a complete native runtime.

The 380-second load time that wasn't

Two observer effects were caught during development. In both cases, a human reviewer reading raw JSON from the Board seat caught something the code path had not made obvious.

The first: a quality smoke reported load_seconds = 380.486. But load_seconds already included materialization, and the report then added materialization again. The real total was 191.829 s. Same data, half the time. A timing field has to say exactly what it includes.

The cold cache that wasn't

The second: a final 7.9 release claimed three "model-cold" replicates with POSIX_FADV_DONTNEED issued, but physical_reads in the output was null, and cold TTFT (12.70 s) was indistinguishable from warm TTFT (12.75 s). The page cache had not been proven cold. We fixed that before release by sampling /proc/<pid>/io from a parent process and recording the SSD and filesystem.

In both cases the code worked. The first report compiled and produced a number. The second run executed and returned ready. The bug was in what the number meant. A verifier has to derive pass/fail from raw counters and refuse to write true when the evidence is incomplete. That is one reason the Board role mattered: someone reading the output cold, from a different seat, was willing to question a result everyone else had already accepted.

8. Try it

From a Linux x86_64 host with AVX-512 VNNI:

python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e '.[runtime]'
export SPARSEFLOW_NATIVE_CACHE="$PWD/.cache/native/int8_vnni"

MODEL=/data/model/Qwen3.6-35B-A3B
INT8=/data/cache/qwen36-int8

PYTHONPATH=src python -m sparseflow doctor "$MODEL" \
  --preset low-memory --int8-container "$INT8" --check-native
PYTHONPATH=src python -m sparseflow prepare-int8 "$MODEL" --output "$INT8"
PYTHONPATH=src python -m sparseflow run "$MODEL" \
  --preset low-memory --int8-container "$INT8" \
  --prompt "Explain sparse expert routing in one paragraph." \
  --max-new-tokens 32

On a 16 GB Windows laptop with AVX-512 VNNI, the same path works with the experimental laptop-16gb preset. The one-shot scripts/setup_windows.ps1 bootstraps uv, Python, Torch, Transformers, and the native extension. Put the project, model, and native build cache on the E: drive if you need to keep C: free.

There is also a local OpenAI-compatible server and a React/Vite frontend. The server is sparseflow serve; the frontend lives under frontend/ and can run against either a real server or fixture mode for development. The server contract is OpenAI-shaped (/v1/chat/completions with SSE), greedy-only, single-request, and stateless, matching the boundaries above.


SparseFlow is a research backend, not a production server. Code, raw result JSON, and result documents are at github.com/GoDiao/SparseFlow. Every performance and correctness claim in this article traces to a specific file under docs/results/ with a clean commit and runtime identity recorded.

Acknowledgments: the tiered-memory streaming design was inspired by JustVugg/colibri (Apache-2.0). The model-independent core, the Qwen3.6 fused-expert layout, the INT8 container format, the AVX-512 VNNI kernel, and the measured results are original to SparseFlow.