Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Run a 35B MoE in About 3 GiB of Active Memory with Edge0-35B-A3B

Edge0-35B-A3B streams experts from SSD so a 4-bit Qwen3.6-35B-A3B derivative decodes at about 15 tok/s in under 3 GiB of active memory.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Edge0-35B-A3B-preview is a 4-bit build of Qwen3.6-35B-A3B from the Edge0 team. It ships with a LoRA adapter and a “prerouter” adapter, both trained for the open-source edge0 streaming inference framework. Its main selling point is memory use. The full checkpoint stays on disk, and expert weights stream from the SSD only when the router picks them. That keeps peak active memory under 3 GiB while decode runs at about 15 tok/s, so a 35B-class sparse MoE fits in phone-class memory. The model is a preview. The card says it is not yet tuned for agent work, and the published speed numbers come only from Apple Silicon.

Key specs

Base model Qwen3.6-35B-A3B
Quantization 4-bit (int4), plus unmerged LoRA adapters and a trained prerouter head
Layers / hidden size 40 / 2048
Experts / active per token 256 / 4
Decode speed 14.9–17.7 tok/s (Mac mini M4 Pro, 24 GB)
Prefill (cold / warm) 113 / 140 tok/s
Peak active memory 2.9 GiB (short contexts; KV cache adds more)
Engines Native, open source, for iOS, macOS, Android, Windows

Three mechanisms make this possible:

  • SSD expert offload. Experts load only when routed to, so peak memory depends on the active set instead of the total parameter count.
  • Prerouter. A trained head predicts the next step’s expert routing, so expert loads run alongside the forward pass. The card reports up to +59% decode throughput, with bigger gains on slower storage.
  • Recover-LoRA. The int4 base is frozen and the LoRA adapters are distilled from the FP teacher. The adapters stay unmerged, so one read-only base can serve several adapter sets.

The card’s quality numbers come from OpenCompass, with the same settings for both models:

Benchmark edge0-35b (int4) Qwen3.6-35B-A3B (fp16)
AIME 2026 86.6 92.7
HumanEval 90.9 95.1
GPQA-Diamond 79.8 81.8
MMLU-Pro 81.0 84.6
IFBench 57.9 61.7
Average 79.2 83.2

The average gap is 3.9 points, but it varies by task: AIME drops 6.1 points, while GPQA-Diamond drops only 2.0.

Install

The card installs edge0 straight from GitHub with the fetch extra, then downloads the model directory:

Terminal window
pip install -e 'git+https://github.com/Edge0-AI/edge0.git#egg=edge0[fetch]'
huggingface-cli download Edge0/Edge0-35b-a3b-preview --local-dir ./Edge0-35b-a3b-preview

You don’t need to wire up the adapters. lora_edge0_35b.safetensors and prerouter_edge0_35b.safetensors sit next to the base checkpoint, and the card says they load automatically.

Run it

Point EDGE0_35B_MODEL at the downloaded directory, then use the edge0 CLI:

Terminal window
export EDGE0_35B_MODEL=$PWD/Edge0-35b-a3b-preview
edge0 chat --name edge0-35b --prompt "Introduce yourself"

The card mentions a Python API and streaming options but doesn’t show them. It points to the edge0 documentation instead. To stick to what the card documents, the examples below wrap the CLI.

The card shows this command only once. It doesn’t say whether edge0 chat --prompt exits after one reply or opens an interactive session. The scripts below pass stdin=subprocess.DEVNULL so they won’t wait on terminal input, but how the command behaves is unverified.

Running several documents through the CLI

This script runs a folder of documents through edge0 chat, one call per file, and saves each output. Note that the card’s “batch serving” use case means something else: one read-only base serving many LoRA adapter sets. This script is just a sequential loop.

batch_summarize.py
import os
import pathlib
import subprocess
# Paths are relative to the current working directory.
# Run this from the folder that contains the model download and docs/.
MODEL_DIR = pathlib.Path("Edge0-35b-a3b-preview").resolve()
INPUT_DIR = pathlib.Path("docs")
OUTPUT_DIR = pathlib.Path("summaries")
OUTPUT_DIR.mkdir(exist_ok=True)
env = dict(os.environ, EDGE0_35B_MODEL=str(MODEL_DIR))
for path in sorted(INPUT_DIR.glob("*.txt")):
text = path.read_text(encoding="utf-8")
prompt = (
"Summarize the following document in five bullet points. "
"Keep names, dates and numbers exact.\n\n" + text
)
result = subprocess.run(
["edge0", "chat", "--name", "edge0-35b", "--prompt", prompt],
capture_output=True,
text=True,
check=True,
env=env,
stdin=subprocess.DEVNULL,
)
(OUTPUT_DIR / f"{path.stem}.md").write_text(result.stdout, encoding="utf-8")
print(f"done: {path.name}")

Some caveats:

  • Start-up cost. The script starts a new edge0 chat process for every file. The card doesn’t report model load or start-up time, so total run time may be much longer than the 15 tok/s decode figure suggests. For many documents, a long-running edge0 serve process may work better (see below), though the card doesn’t document its API.
  • Argument length. Each full document goes into a single command-line argument. Large files can hit OS limits (OSError: Argument list too long), and the card documents no way to pass a prompt through stdin or a file.
  • Context size. Keep inputs short. Every extra token of context grows the KV cache, and the 2.9 GiB figure only holds for short contexts.
  • Output format. The card doesn’t describe what edge0 chat prints to stdout, so check one output before you rely on the format.

Private extraction from a local file

The card targets edge devices where GPU memory is scarce and storage is fast. One example is pulling structured fields out of a file that should never leave the machine:

extract_invoice.py
import os
import pathlib
import subprocess
import sys
# Resolved relative to the current working directory.
MODEL_DIR = pathlib.Path("Edge0-35b-a3b-preview").resolve()
env = dict(os.environ, EDGE0_35B_MODEL=str(MODEL_DIR))
source = pathlib.Path(sys.argv[1]).read_text(encoding="utf-8")
prompt = f"""Extract these fields from the invoice below and answer with JSON only:
vendor_name, invoice_number, invoice_date, total_amount, currency.
Use null for any field that is missing.
Invoice:
{source}
"""
result = subprocess.run(
["edge0", "chat", "--name", "edge0-35b", "--prompt", prompt],
capture_output=True,
text=True,
check=True,
env=env,
stdin=subprocess.DEVNULL,
)
print(result.stdout)

Run it with python extract_invoice.py invoice.txt from the folder that contains the model download. The card says the bundled chat template enables thinking mode. It doesn’t say whether edge0 chat prints the reasoning text, so stdout may contain more than the JSON. Parse it defensively instead of passing stdout straight to json.loads. The model scores 57.9 on IFBench (instruction following), so also check that the output matches your schema. The argument-length caveat above applies here too.

A local OpenAI-compatible endpoint

If you’d rather call the model over HTTP from an existing app, edge0 can serve an OpenAI-compatible API:

serve.sh
#!/usr/bin/env bash
set -euo pipefail
export EDGE0_35B_MODEL="$PWD/Edge0-35b-a3b-preview"
edge0 serve --name edge0-35b --port 8085

The card stops at this command. It doesn’t list the route paths, the model identifier the server expects, or which OpenAI parameters work, so read the edge0 docs before you point a client at port 8085. The card also says one read-only base can serve many LoRA adapter sets, but it doesn’t show how to load or switch adapters when serving.

Gotchas

  • It’s a preview. The card says coverage and quality are still being extended and that tuning mainly targets the base model’s languages.
  • Weak at agent tasks. The card says tool use, multi-step planning and long-horizon autonomy are “currently weak” in this release. Don’t build an agent loop on it yet.
  • The 3 GiB figure is peak active memory at short contexts. The KV cache grows with context length, so long prompts raise peak memory. The card advises keeping contexts short if you want to stay near 3 GiB. It doesn’t report total system RAM use, including any OS page cache used by streamed experts.
  • Storage speed matters. Experts stream from disk on demand, so the design assumes fast NVMe or internal flash. On a slow disk, expect lower throughput, even though the prerouter’s gain grows with storage latency.
  • Speed numbers come from one Mac so far. All performance figures come from the MLX backend on a Mac mini M4 Pro with 24 GB. The card says the iOS, Android and Windows engines are “still maturing”.
  • Expect some quality loss from quantization. It averages 3.9 points against fp16 and is largest on AIME (−6.1) and HumanEval (−4.2).
  • Keep the directory as downloaded. The card says the adapters are co-located with the base checkpoint and load automatically. It doesn’t say what happens if you move them, so the safe choice is to leave the directory as is.
  • Repo name casing. The repo id is Edge0/Edge0-35B-A3B-preview, but the card’s download command uses Edge0/Edge0-35b-a3b-preview. If the download fails, try the other casing.
  • Install comes from git. edge0 installs as an editable package from the GitHub repo, not from a pinned PyPI release. Pin a commit if you need reproducible builds.

When to pick it

Pick Edge0-35B-A3B when you want a 35B-class MoE on hardware that can’t hold one in memory, for offline reasoning, code generation or chat. The engines target iOS, macOS, Android and Windows, but the only published speed figure is about 15 tok/s decode on a Mac mini M4 Pro with 24 GB, which the card calls fast enough for interactive use. On average the 4-bit pipeline scores 3.9 points below the fp16 base. The gap is 2.0 on GPQA-Diamond and as much as 6.1 on AIME math. The Apache 2.0 license allows commercial use.

Skip it in these cases:

  • You need tool calling or multi-step agents. The card says these are weak in this preview.
  • You have enough GPU memory for the fp16 base. Running Qwen3.6-35B-A3B directly avoids the quality gap.
  • You need long-context work at low memory. The KV cache will push you past 3 GiB.
  • You need proven performance on phones, Windows or Android today. No numbers are published for those yet.
  • You need a stable, documented Python API. The card points to the edge0 docs for that, and this is an early preview.

Related