Model deep dive·
Run a 35B MoE in About 3 GiB of Active Memory with Edge0-35B-A3B
Edge0-35B-A3B streams experts from SSD so a 4-bit Qwen3.6-35B-A3B derivative decodes at about 15 tok/s in under 3 GiB of active memory.
Edge0/Edge0-35B-A3B-preview
Edge0-35B-A3B streams experts from SSD so a 4-bit Qwen3.6-35B-A3B derivative decodes at about 15 tok/s in under 3 GiB of active memory.
Edge0/Edge0-35B-A3B-preview
MiniCPM5-2B is a 2.5B Llama-architecture model with 131K context and tool calling. Here's how to serve it and what it's good for.
openbmb/MiniCPM5-2B