Run a 35B MoE in About 3 GiB of Active Memory with Edge0-35B-A3B
Edge0-35B-A3B streams experts from SSD so a 4-bit Qwen3.6-35B-A3B derivative decodes at about 15 tok/s in under 3 GiB of active memory.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Edge0-35B-A3B-preview is a 4-bit build of Qwen3.6-35B-A3B from the Edge0 team. It ships with a LoRA adapter and a “prerouter” adapter, both trained for the open-source edge0 streaming inference framework. Its main selling point is memory use. The full checkpoint stays on disk, and expert weights stream from the SSD only when the router picks them. That keeps peak active memory under 3 GiB while decode runs at about 15 tok/s, so a 35B-class sparse MoE fits in phone-class memory. The model is a preview. The card says it is not yet tuned for agent work, and the published speed numbers come only from Apple Silicon.
Key specs
| Base model | Qwen3.6-35B-A3B |
| Quantization | 4-bit (int4), plus unmerged LoRA adapters and a trained prerouter head |
| Layers / hidden size | 40 / 2048 |
| Experts / active per token | 256 / 4 |
| Decode speed | 14.9–17.7 tok/s (Mac mini M4 Pro, 24 GB) |
| Prefill (cold / warm) | 113 / 140 tok/s |
| Peak active memory | 2.9 GiB (short contexts; KV cache adds more) |
| Engines | Native, open source, for iOS, macOS, Android, Windows |
Three mechanisms make this possible:
- SSD expert offload. Experts load only when routed to, so peak memory depends on the active set instead of the total parameter count.
- Prerouter. A trained head predicts the next step’s expert routing, so expert loads run alongside the forward pass. The card reports up to +59% decode throughput, with bigger gains on slower storage.
- Recover-LoRA. The int4 base is frozen and the LoRA adapters are distilled from the FP teacher. The adapters stay unmerged, so one read-only base can serve several adapter sets.
The card’s quality numbers come from OpenCompass, with the same settings for both models:
| Benchmark | edge0-35b (int4) | Qwen3.6-35B-A3B (fp16) |
|---|---|---|
| AIME 2026 | 86.6 | 92.7 |
| HumanEval | 90.9 | 95.1 |
| GPQA-Diamond | 79.8 | 81.8 |
| MMLU-Pro | 81.0 | 84.6 |
| IFBench | 57.9 | 61.7 |
| Average | 79.2 | 83.2 |
The average gap is 3.9 points, but it varies by task: AIME drops 6.1 points, while GPQA-Diamond drops only 2.0.
Install
The card installs edge0 straight from GitHub with the fetch extra, then downloads the model directory:
pip install -e 'git+https://github.com/Edge0-AI/edge0.git#egg=edge0[fetch]'
huggingface-cli download Edge0/Edge0-35b-a3b-preview --local-dir ./Edge0-35b-a3b-previewYou don’t need to wire up the adapters. lora_edge0_35b.safetensors and prerouter_edge0_35b.safetensors sit next to the base checkpoint, and the card says they load automatically.
Run it
Point EDGE0_35B_MODEL at the downloaded directory, then use the edge0 CLI:
export EDGE0_35B_MODEL=$PWD/Edge0-35b-a3b-previewedge0 chat --name edge0-35b --prompt "Introduce yourself"The card mentions a Python API and streaming options but doesn’t show them. It points to the edge0 documentation instead. To stick to what the card documents, the examples below wrap the CLI.
The card shows this command only once. It doesn’t say whether edge0 chat --prompt exits after one reply or opens an interactive session. The scripts below pass stdin=subprocess.DEVNULL so they won’t wait on terminal input, but how the command behaves is unverified.
Running several documents through the CLI
This script runs a folder of documents through edge0 chat, one call per file, and saves each output. Note that the card’s “batch serving” use case means something else: one read-only base serving many LoRA adapter sets. This script is just a sequential loop.
import osimport pathlibimport subprocess
# Paths are relative to the current working directory.# Run this from the folder that contains the model download and docs/.MODEL_DIR = pathlib.Path("Edge0-35b-a3b-preview").resolve()INPUT_DIR = pathlib.Path("docs")OUTPUT_DIR = pathlib.Path("summaries")OUTPUT_DIR.mkdir(exist_ok=True)
env = dict(os.environ, EDGE0_35B_MODEL=str(MODEL_DIR))
for path in sorted(INPUT_DIR.glob("*.txt")): text = path.read_text(encoding="utf-8") prompt = ( "Summarize the following document in five bullet points. " "Keep names, dates and numbers exact.\n\n" + text ) result = subprocess.run( ["edge0", "chat", "--name", "edge0-35b", "--prompt", prompt], capture_output=True, text=True, check=True, env=env, stdin=subprocess.DEVNULL, ) (OUTPUT_DIR / f"{path.stem}.md").write_text(result.stdout, encoding="utf-8") print(f"done: {path.name}")Some caveats:
- Start-up cost. The script starts a new
edge0 chatprocess for every file. The card doesn’t report model load or start-up time, so total run time may be much longer than the 15 tok/s decode figure suggests. For many documents, a long-runningedge0 serveprocess may work better (see below), though the card doesn’t document its API. - Argument length. Each full document goes into a single command-line argument. Large files can hit OS limits (
OSError: Argument list too long), and the card documents no way to pass a prompt through stdin or a file. - Context size. Keep inputs short. Every extra token of context grows the KV cache, and the 2.9 GiB figure only holds for short contexts.
- Output format. The card doesn’t describe what
edge0 chatprints to stdout, so check one output before you rely on the format.
Private extraction from a local file
The card targets edge devices where GPU memory is scarce and storage is fast. One example is pulling structured fields out of a file that should never leave the machine:
import osimport pathlibimport subprocessimport sys
# Resolved relative to the current working directory.MODEL_DIR = pathlib.Path("Edge0-35b-a3b-preview").resolve()env = dict(os.environ, EDGE0_35B_MODEL=str(MODEL_DIR))
source = pathlib.Path(sys.argv[1]).read_text(encoding="utf-8")
prompt = f"""Extract these fields from the invoice below and answer with JSON only:vendor_name, invoice_number, invoice_date, total_amount, currency.Use null for any field that is missing.
Invoice:{source}"""
result = subprocess.run( ["edge0", "chat", "--name", "edge0-35b", "--prompt", prompt], capture_output=True, text=True, check=True, env=env, stdin=subprocess.DEVNULL,)print(result.stdout)Run it with python extract_invoice.py invoice.txt from the folder that contains the model download. The card says the bundled chat template enables thinking mode. It doesn’t say whether edge0 chat prints the reasoning text, so stdout may contain more than the JSON. Parse it defensively instead of passing stdout straight to json.loads. The model scores 57.9 on IFBench (instruction following), so also check that the output matches your schema. The argument-length caveat above applies here too.
A local OpenAI-compatible endpoint
If you’d rather call the model over HTTP from an existing app, edge0 can serve an OpenAI-compatible API:
#!/usr/bin/env bashset -euo pipefail
export EDGE0_35B_MODEL="$PWD/Edge0-35b-a3b-preview"edge0 serve --name edge0-35b --port 8085The card stops at this command. It doesn’t list the route paths, the model identifier the server expects, or which OpenAI parameters work, so read the edge0 docs before you point a client at port 8085. The card also says one read-only base can serve many LoRA adapter sets, but it doesn’t show how to load or switch adapters when serving.
Gotchas
- It’s a preview. The card says coverage and quality are still being extended and that tuning mainly targets the base model’s languages.
- Weak at agent tasks. The card says tool use, multi-step planning and long-horizon autonomy are “currently weak” in this release. Don’t build an agent loop on it yet.
- The 3 GiB figure is peak active memory at short contexts. The KV cache grows with context length, so long prompts raise peak memory. The card advises keeping contexts short if you want to stay near 3 GiB. It doesn’t report total system RAM use, including any OS page cache used by streamed experts.
- Storage speed matters. Experts stream from disk on demand, so the design assumes fast NVMe or internal flash. On a slow disk, expect lower throughput, even though the prerouter’s gain grows with storage latency.
- Speed numbers come from one Mac so far. All performance figures come from the MLX backend on a Mac mini M4 Pro with 24 GB. The card says the iOS, Android and Windows engines are “still maturing”.
- Expect some quality loss from quantization. It averages 3.9 points against fp16 and is largest on AIME (−6.1) and HumanEval (−4.2).
- Keep the directory as downloaded. The card says the adapters are co-located with the base checkpoint and load automatically. It doesn’t say what happens if you move them, so the safe choice is to leave the directory as is.
- Repo name casing. The repo id is
Edge0/Edge0-35B-A3B-preview, but the card’s download command usesEdge0/Edge0-35b-a3b-preview. If the download fails, try the other casing. - Install comes from git. edge0 installs as an editable package from the GitHub repo, not from a pinned PyPI release. Pin a commit if you need reproducible builds.
When to pick it
Pick Edge0-35B-A3B when you want a 35B-class MoE on hardware that can’t hold one in memory, for offline reasoning, code generation or chat. The engines target iOS, macOS, Android and Windows, but the only published speed figure is about 15 tok/s decode on a Mac mini M4 Pro with 24 GB, which the card calls fast enough for interactive use. On average the 4-bit pipeline scores 3.9 points below the fp16 base. The gap is 2.0 on GPQA-Diamond and as much as 6.1 on AIME math. The Apache 2.0 license allows commercial use.
Skip it in these cases:
- You need tool calling or multi-step agents. The card says these are weak in this preview.
- You have enough GPU memory for the fp16 base. Running Qwen3.6-35B-A3B directly avoids the quality gap.
- You need long-context work at low memory. The KV cache will push you past 3 GiB.
- You need proven performance on phones, Windows or Android today. No numbers are published for those yet.
- You need a stable, documented Python API. The card points to the edge0 docs for that, and this is an early preview.
Related
DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
GLM-5.3-Flash: What You Can Run With Z.ai's 18B-Active Multimodal Model
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
Run K2-Horizon-MoVA-36B-A4B for Agents and 512K-Token Context
IFM's open MoE model runs 4B active parameters with a 512K context window. How to serve it, call it, and use it for agents and long documents.
IFM/K2-Horizon-MoVA-36B-A4B