July 13 → July 23 · a 16 GB laptop → the same 16 GB laptop
Running a 67 GB MoE on a 16 GB Laptop
A ten-day build log: what we tried, what the numbers said, where our assumptions broke, and how we actually got Qwen3.6-35B-A3B running on the same 16 GB Lenovo we started on.
Where this starts
On July 13, I opened a Codex session on my laptop. It’s a pretty ordinary machine: 16 GB of RAM, a 6-core CPU, and no GPU that could usefully run this model. I started by feeding it a GitHub link to JustVugg/colibri—a project that somehow streams a massive 744B MoE directly from disk. I wanted to understand how they pulled that off, and more importantly, whether the same idea could run a smaller MoE on my own hardware.
We picked Qwen/Qwen3.6-35B-A3B as our target. It is a large model: 35B total parameters, about 3B active per token, 40 layers, 256 routed experts per layer, top-8 routing, one shared expert, Gated DeltaNet, Gated Attention, plus vision and MTP components. The full BF16 checkpoint is 67 GiB. A direct from_pretrained on a 16 GB laptop would run out of memory before the load completed.
To build this efficiently, I set up a "two-seat" workflow. I used a rented Xeon box (125 GiB RAM, fast NVMe) as my Implementation Seat to write the runtime and run heavy experiments. Meanwhile, my local laptop acted as the Review Board: I pulled clean commits here, parsed raw JSON, and carefully checked performance claims before merging them. This separation made the constraint visible early: if a feature worked on the big Xeon box but choked on the laptop's limited hardware, it did not pass.
1. The premise
If you inspect the Safetensors header—a small metadata block whose length is stored in the first 8 bytes—you can map the model without loading the payload. Once the tensors are grouped by responsibility, the memory problem becomes surprisingly concrete:
Three numbers from this map became the core constraints of the project:
- 6.00 MiB: The size of a single routed expert. It's just a byte range inside a fused
gate_up_proj/down_projtensor, not an individual file. - 1.88 GiB: The theoretical worst-case data we have to read from disk *per token* if we have zero cache hits. This is the bottleneck we are fighting.
- 6.97 GiB: The dense resident weights. This is the absolute bare minimum that must sit in RAM at all times.
2. The architecture
We settled on a three-layer architecture with one golden rule: the core engine must be completely decoupled from Qwen. The host runtime handles attention, Gated DeltaNet, KV states, tokenization, and generation loops. SparseFlow's job is strictly focused: expert lookup, hot cache management, streaming I/O, and running our native W8A8 kernel. They talk through a very clean, narrow interface.
ExpertLocator: An expert is just a byte range
Qwen3.6 stores its routed experts in a fused layout. Expert 17 at Layer 0 isn't a standalone file; it's a specific slice along axis 0 of a giant fused tensor. The locator's job is to calculate the exact byte offset, length, dtype, and shape from the metadata header, and validate everything (expert counts, slice bounds, etc.) before we touch the disk. On the real model, Layer 0 Expert 17 resolves to exactly 6,291,456 bytes (6.00 MiB) across two shards, which we verified via SHA-256.
ExpertCache: Leases and byte budgets
The cache is a per-layer LRU with an optional global byte budget. The trickiest part here was implementing a lease system: when a forward pass pins an expert, the cache is forbidden from evicting it until the forward pass is done. If the cache is completely full of leased experts and we hit another cache miss, the provider falls back to returning a transient payload instead of corrupting an active tensor view.
finally block, and assert that the active lease count is exactly zero at the end of every token generation.
INT8 container: canonical-int8-v1
We packed everything into one .sfi file per layer: expert-major row-major S8, per-output-channel symmetric quantization, FP16 scales, 4 KiB alignment, and checksums everywhere. The conversion was processed one expert at a time, meaning we never had to materialize a giant fused layer in memory:
| Metric | Value |
|---|---|
| Routed BF16 source | 60.000 GiB |
| INT8 physical storage | 30.078 GiB (50.13%) |
| Experts / layers | 10,240 / 40 |
| Conversion wall time | 373.03 s |
| Conversion peak RSS | 630.84 MiB |
| Max abs dequant error (samples) | 0.00201 |
We also precomputed the INT32 row sums into a sidecar file canonical-int8-exec-v1 (125.99 MiB), making runtime row-sum preparation practically instant.
3. Correctness first
Before optimizing for speed, we set a strict rule: the resident path (all in RAM) and the streaming path (from disk) must produce bit-identical outputs. No "close enough", no epsilon tolerances. We compared the full 248,320-way vocabulary logits at every single step (both prefill and decode) using SHA-256 hashes.
max_abs_error = 0, I finally stopped worrying about correctness and started focusing on speed.
To prove this, we ran a 60-question evaluation set (HellaSwag, ARC-Challenge, MMLU, 20 questions each, seed=1234) comparing BF16 resident, W8A8 native resident, and W8A8 native streaming with a 4 GiB cache. The native resident and streaming paths matched perfectly on every single prediction, choice, and token log-likelihood. The streaming path read 767.36 GiB of expert payload over the entire run while strictly respecting the 4 GiB cache limit—and gave the exact same answers.
Formal matrix · Stage 7.5.6 · Xeon Gold 6248R · 10 threads
4. The numbers (on the big box)
Formal Stage 7.5.6 protocol: Xeon Gold 6248R, 10 CPU threads (we tested 1, 5, 10, and 20; 10 threads gave 2.934 tok/s, while 20 threads actually degraded to 2.837 due to overhead), one fixed prompt, 32 greedy tokens, one warmup run, and three measured runs. For the 4 GiB streaming path, we used three independent, cold-started processes to ensure no OS page-caching was cheating the results.
| Path | Decode | TTFT | RSS | Read/token |
|---|---|---|---|---|
| BF16 resident (ceiling) | 3.43 tok/s | 4.67 s | 65.5 GiB | 0 |
| W8A8 native resident | 2.49 tok/s | 8.25 s | 35.8 GiB | 0 |
| W8A8 streaming 4 GiB (warm) | 1.41 tok/s | 12.36 s | 9.87 GiB | 355 MiB |
| W8A8 streaming 4 GiB (cold) | 1.02 tok/s | 57.1 s | 9.75 GiB | 355 MiB |
| Generic offload (warm) | 0.028 tok/s | 38.6 s | 5.8 GiB | n/a |
The 9.2 GiB moment
Stage 7.1, July 15. A fresh process generated four tokens from the 67 GiB model and peaked at 9.216 GiB RSS. The generated text was "Here's a thinking". The process never allocated the 60 GiB of expert tensors, not even temporarily.
9.216 GiB number for a while. Before
Stage 7.1, "streaming experts" was a plausible design with good math behind it.
After that run, it was a real runtime that fit in roughly a quarter of the
model's size. The caveat mattered too: Linux page cache is outside process RSS,
so we did not claim it as resident model memory. The process-level number was
real and reproducible. The question had shifted from "can this work?" to "where
does it break?"
5. Where it broke (and what we learned from stopping)
The negative results turned out to be some of the most useful results in the project. Each one was measured, gated, and kept out of the default path.
| Optimization | Hypothesis | Measured | Verdict |
|---|---|---|---|
| Pure fused decode kernel | Faster expert micro-benchmark means faster end-to-end decode | 0.886× legacy end-to-end | NO-GO as default |
| Shared streaming batching | Route union at B=8 should save shared-cache reads | B=8 / 4 GiB union reads 1.205× more than round-robin | NO-GO |
| lm_head native rewrite | lm_head looks large enough to justify a rewrite | About 5% of decode | NO-GO (below 10% gate) |
| Deterministic current-layer I/O pipeline | Overlap expert I/O with shared-expert compute | 0.90 vs 1.21 tok/s synchronous | NO-GO as default |
| DeltaNet projection fusion | Small projection fusion should remove a visible decode cost | Less than 1% of total decode | NO-GO |
The micro-benchmark trap
The pure fused decode kernel was the instructive failure. Its isolated eight-expert micro-benchmark was faster than the canonical per-expert path. Paired same-process AB/BA tests on the full runtime measured the fused path at 0.886× the legacy decode. It reduced routed-MoE time, but changed CPU frequency, thread-pool behavior, and memory access patterns enough to lose end-to-end throughput.
When batching makes things worse
Stage 7.7 tested whether multiple concurrent requests could share a streaming cache and amortize reads. The route overlap was real: at B=8, the union of selected experts was 1.749× smaller than the sum of independent selections. The grouped kernel was exact and reached a 1.707× aggregate speedup at B=4. The failure was in the shared-cache gate:
| Batch / Cache | Union hit rate | Round-robin hit rate | Union / RR reads |
|---|---|---|---|
| B=4 / 4 GiB | 57.7% | 72.5% | 1.012× |
| B=8 / 4 GiB | 8.1% | 56.4% | 1.205× |
| B=4 / 8 GiB | 78.9% | 86.1% | 1.00× |
| B=8 / 8 GiB | 67.4% | 81.3% | 1.00× |
B=8 / 4 GiB was the failure: 122.4 GiB union versus 101.6 GiB round-robin. At 8 GiB the result was close to neutral, too weak to ship a scheduler. The decision was explicit: resident-only fixed-cohort batching is implemented; shared streaming batching is disabled and labeled a known limitation.
6. The laptop closes the loop
On July 23, ten days after the project started on the same machine, I closed
most other applications on the 16 GB Lenovo, ran doctor with the
real INT8 container, and got ready=true. Available RAM was
9.95 GiB, the minimum was 8.86 GiB, and the headroom was
1.08 GiB. The status was warn. Then came a two-token generation:
256 MiB expert cache
model load: 7.17 s
prefill: 21.83 s
single decode: 2.23 s (~0.45 tok/s)
generated: "Here's"
expert reads: ~7.77 GiB
The right framing is: experimental laptop preset, 0.45 tok/s, about 1 GiB headroom, and other applications closed. It is a proof of concept. The point is not that a 16 GB laptop is suddenly a comfortable 35B inference machine; the point is that the memory boundary is no longer an absolute wall.
7. Boundaries, and two bugs caught late
Here is what the Public Alpha is, and is not:
| Path | Status |
|---|---|
| INT8 native resident hybrid | Stable baseline |
| INT8 native single-request streaming S1 LRU | Stable low-memory |
| INT8 native resident grouped fixed cohort | Experimental opt-in |
| Shared streaming batching / subcohort | Disabled, known limitation |
| Pure fused decode | Diagnostics only |
Generation is greedy-only. Conversation is stateless: the server replays the full message list rather than persisting KV state. AVX-512 VNNI is required. The host runtime still owns attention, Gated DeltaNet, KV state, tokenization, and the generation loop. SparseFlow is not yet a complete native runtime.
The 380-second load time that wasn't
Two observer effects were caught during development. In both cases, a human reviewer reading raw JSON from the Board seat caught something the code path had not made obvious.
The first: a quality smoke reported load_seconds = 380.486.
But load_seconds already included materialization, and the report
then added materialization again. The real total was 191.829 s. Same data,
half the time. A timing field has to say exactly what it includes.
The cold cache that wasn't
The second: a final 7.9 release claimed three "model-cold" replicates with
POSIX_FADV_DONTNEED issued, but physical_reads in the
output was null, and cold TTFT (12.70 s) was indistinguishable
from warm TTFT (12.75 s). The page cache had not been proven cold. We fixed
that before release by sampling /proc/<pid>/io from a parent
process and recording the SSD and filesystem.
ready. The bug was in what the
number meant. A verifier has to derive pass/fail from raw counters and refuse to
write true when the evidence is incomplete. That is one reason the
Board role mattered: someone reading the output cold, from a different seat,
was willing to question a result everyone else had already accepted.
8. Try it
From a Linux x86_64 host with AVX-512 VNNI:
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e '.[runtime]'
export SPARSEFLOW_NATIVE_CACHE="$PWD/.cache/native/int8_vnni"
MODEL=/data/model/Qwen3.6-35B-A3B
INT8=/data/cache/qwen36-int8
PYTHONPATH=src python -m sparseflow doctor "$MODEL" \
--preset low-memory --int8-container "$INT8" --check-native
PYTHONPATH=src python -m sparseflow prepare-int8 "$MODEL" --output "$INT8"
PYTHONPATH=src python -m sparseflow run "$MODEL" \
--preset low-memory --int8-container "$INT8" \
--prompt "Explain sparse expert routing in one paragraph." \
--max-new-tokens 32
On a 16 GB Windows laptop with AVX-512 VNNI, the same path works with
the experimental laptop-16gb preset. The one-shot
scripts/setup_windows.ps1 bootstraps uv, Python,
Torch, Transformers, and the native extension. Put the project, model, and
native build cache on the E: drive if you need to keep C: free.
There is also a local OpenAI-compatible server and a React/Vite frontend.
The server is sparseflow serve; the frontend lives under
frontend/ and can run against either a real server or fixture mode
for development. The server contract is OpenAI-shaped
(/v1/chat/completions with SSE), greedy-only, single-request, and
stateless, matching the boundaries above.
SparseFlow is a research backend, not a production server.
Code, raw result JSON, and result documents are at
github.com/GoDiao/SparseFlow.
Every performance and correctness claim in this article traces to a specific
file under docs/results/ with a clean commit and runtime identity
recorded.
Acknowledgments: the tiered-memory streaming design was inspired by JustVugg/colibri (Apache-2.0). The model-independent core, the Qwen3.6 fused-expert layout, the INT8 container format, the AVX-512 VNNI kernel, and the measured results are original to SparseFlow.