Run 95.5 GiB Qwen3.8-Flash-Next on a single 64 GB Mac at 41–52 tok/s.
Slipstream is a lean, high-performance C++ and Metal inference engine built specifically for Apple Silicon. It combines SSD expert streaming with predictive read-ahead and Prompt Lookup + MTP speculative drafting to serve frontier-scale models that exceed your Mac's physical RAM.
It serves Qwen3.8-Flash-Next V3 (125.7B parameters, 512 routed experts, 7.3B active per token) and its Swift KV-sparse variant at 1.76x the speed of llama.cpp, with context scaling tested all the way out to 130,000 tokens without decode collapse.
Everything is open source under Apache-2.0.
| Model Variant | HF Checkpoint (GGUF) | Active / Total Weights | Reasoning (Scorecard) | Peak Speed | RAM Needed |
|---|---|---|---|---|---|
| Swift-Flash-Next V3 (New KV-Sparse) | nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF | 7.3B / 125.7B | 70.3% (GPQA 54.3%, MATH 62.9%) | 41–52 tok/s | 64 GB Mac |
| Qwen3.8-Flash-Next V3 (34k+ Downloads) | nitinpanj/qwen38-flash-next-v3 | 7.3B / 125.7B | 67.6% (GPQA 45.7%, MATH 60.0%) | 40–48 tok/s | 64 GB Mac |
- Frontier Reasoning on a Laptop: Outperforms standard 27B dense models on GPQA Diamond (54.3% vs 45.7%) and MATH-500 (62.9% vs 60.0%), with 92% on HumanEval and 96% on GSM8K.
- Fast Generation: 41–52 tok/s sustained decode on an M5 Pro (64 GB).
-
No Context Collapse: Maintains 32–43 tok/s out to 130,000 tokens thanks to 48 recurrent linear DeltaNet layers (
$O(1)$ state growth) and only 16 full-attention layers. - KV-Sparsity (Swift): The Swift variant incorporates KV-sparse attention distilled from Swift-1.5 with spliced Q8 donor backbones, minimizing RAM growth in long agent sessions.
- Hardware: Apple Silicon Mac with 64 GB Unified Memory (M2/M3/M4/M5 Pro/Max).
- Disk: ~100 GB for the multi-shard GGUF weights + ~95 GB working disk space.
- macOS: macOS 15.0+ with Command Line Tools or Xcode installed (
xcode-select --install).
git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4Note: make compiles the native C++ runtime and Metal compute kernels into build/slipstream and build/slipstream.metallib in under a minute.
We recommend the Swift KV-sparse variant for optimal reasoning accuracy and lower KV memory footprint:
# Recommended: Swift-Qwen3.8-Flash-Next V3 (95.5 GiB)
huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3
# Or download with fast parallel transfer if hf_transfer is installed:
HF_HUB_ENABLE_HF_TRANSFER=1 huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3(Alternative: If you prefer the plain dense base model without Swift KV-sparsity:)
huggingface-cli download nitinpanj/qwen38-flash-next-v3 \
--local-dir ~/models/qwen38-flash-next-v3On 64 GB Macs, macOS defaults the wired GPU limit to ~48 GiB. Raise it to 58 GiB so the SSD expert cache and KV pool have ample headroom:
sudo sysctl iogpu.wired_limit_mb=59392Point ./slipstream serve directly at the downloaded model directory:
./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into
<model-dir>/prepared/(~5–7 minutes). Subsequent launches load in ~10–15 seconds.
The server provides a standard OpenAI-compatible API on http://127.0.0.1:8090:
curl -s http://127.0.0.1:8090/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "local/swift-qwen38-flash-next-v3",
"messages": [
{"role": "user", "content": "Write a clean, optimal Python function for interval merging."}
],
"temperature": 0.0
}'from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8090/v1", api_key="not-needed")
response = client.chat.completions.create(
model="local/swift-qwen38-flash-next-v3",
messages=[{"role": "user", "content": "Explain multi-head self-attention with linear algebra."}],
temperature=0.0,
)
print(response.choices[0].message.content)# Launch omp connected directly to Slipstream
omp --model splash-flashnext/local/swift-qwen38-flash-next-v3 \
--tools=read,write,edit,bash,grep,glob,todo \
--thinking=low --approval-mode=yoloEvaluated on Apple MacBook M5 Pro (64 GB Unified Memory, Temperature 0.0):
| Task Domain | Benchmark / Prompt | llama.cpp Fork | Slipstream | Speedup | llama.cpp TTFT | Slipstream TTFT |
|---|---|---|---|---|---|---|
| Math Reasoning | GSM8K (eggs derivation) | 24.0 tok/s | 43.6 tok/s | 1.82x | 4,024 ms | 2,337 ms |
| Math Derivation | MATH-500 series ( |
24.3 tok/s | 43.1 tok/s | 1.77x | 1,655 ms | 1,587 ms |
| Constraint Logic | 3-chair deduction | 25.4 tok/s | 46.0 tok/s | 1.81x | 1,469 ms | 1,042 ms |
| Python Coding |
merge_intervals ($O(N \log N)$) |
19.7 tok/s | 35.0 tok/s | 1.77x | 1,507 ms | 1,070 ms |
| Systems Coding | Rust CSV parser | 22.7 tok/s | 37.5 tok/s | 1.65x | 1,257 ms | 859 ms |
| Tech Communication | Multi-head attention | 22.5 tok/s | 39.4 tok/s | 1.75x | 1,267 ms | 843 ms |
| OVERALL AVERAGE | Across all 6 domains | 23.1 tok/s | 40.8 tok/s | 1.76x | 1,863 ms | 1,290 ms |
Evaluated across 145 standardized items under memory guard (Seed 1234, T=0.0):
| Benchmark | Items | Plain Flash-Next V3 | Swift-Flash-Next V3 | Accuracy Delta |
|---|---|---|---|---|
| AIME 2025 | 20 | 45.0% (9/20) | 45.0% (9/20) | 0.0% |
| MATH-500 (L4–5) | 35 | 60.0% (21/35) | 62.9% (22/35) | +2.9% |
| GPQA Diamond | 35 | 45.7% (16/35) | 54.3% (19/35) | +8.6% |
| GSM8K | 25 | 96.0% (24/25) | 96.0% (24/25) | 0.0% |
| HumanEval | 25 | 92.0% (23/25) | 92.0% (23/25) | 0.0% |
| Hard Logic | 5 | 100.0% (5/5) | 100.0% (5/5) | 0.0% |
| OVERALL SCORECARD | 145 | 67.6% (98/145) | 70.3% (102/145) | +2.8% |
Measured across 3,086 live agent requests on Apple Silicon (M5 Pro 64 GB):
| Context Range (Tokens) | Live Runs | Average Decode | Median Decode (p50) | Peak Decode | Avg TTFT |
|---|---|---|---|---|---|
| < 1,000 | 314 | 41.5 tok/s | 41.9 tok/s | 59.8 tok/s | 2.16 s |
| 1k – 4,000 | 21 | 41.0 tok/s | 42.5 tok/s | 64.5 tok/s | 5.26 s |
| 4k – 8,000 | 58 | 43.6 tok/s | 43.2 tok/s | 67.2 tok/s | 7.36 s |
| 8k – 16,000 | 117 | 43.6 tok/s | 44.6 tok/s | 58.2 tok/s | 7.91 s |
| 16k – 32,000 | 562 | 38.2 tok/s | 40.9 tok/s | 58.0 tok/s | 13.59 s |
| 32k – 64,000 | 1,029 | 35.0 tok/s | 37.5 tok/s | 55.6 tok/s | 13.24 s |
| 64k – 96,000 | 650 | 32.4 tok/s | 34.7 tok/s | 53.9 tok/s | 12.81 s |
| 96k – 130,000 | 364 | 32.9 tok/s | 33.3 tok/s | 43.8 tok/s | 7.95 s |
The core primitives implemented in Slipstream:
- 512-route sparse MoE streaming with predictive read-ahead
- Quasi-Sparse Attention (QSA) indexer and selection kernels
- Hyper-connection mixing and per-layer embedding gathers
- Metal GPU-mapped n-gram tables
- Single-lane speculative verification with Prompt Lookup Decoding (PLD) + Multi-Token Prediction (MTP)
...were engineered to match the upcoming model generation. Assuming Qwen4 follows Flash-Next's architectural blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream can serve as a direct template to run Qwen4 locally on Apple Silicon on day one.
- Current Status: Tested exclusively on an Apple MacBook Pro (M5 Pro, 64 GB Unified Memory).
- Porting: Because I do not have access to modern NVIDIA (CUDA) or AMD (ROCm) GPU hardware, I cannot build and test those backends myself.
- Collaboration Offer: If anyone in the community has hardware available and wants to bring these expert-streaming and speculative decoding gains to CUDA or ROCm, I am happy to collaborate and help with the port. Open an issue or reach out!
- Splash Team (Incoai): Full credit to the creators of Splash (github.com/incoai/splash). Their C++ Metal speculative decoding design and memory architecture provided the foundation for this work. We will prepare a clean PR/patch proposing these Flash-Next and SSD streaming extensions to the Splash upstream repo.
- ds4 Team: For their valuable insights on Metal router numerical precision (Taylor polynomial softplus expansion) and streaming scheduling designs.
- Qwen Team: For training Qwen3.8-Flash-Next and open-sourcing the hybrid linear MTP architecture.
- ukisai: For the Swift-1.5 distillation work enabling KV-sparse reasoning.
- bartowski & unsloth: For donor quants and quantization tooling.
- mihailescu2m: For initial expert streaming concepts in llama.cpp.
Apache-2.0. See LICENSE.


