Model deep dive · Oct 8, 2026
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
Edge0/Audio8-ASR-Infinite
Model deep dive · Oct 8, 2026
Tencent's MIT-licensed AuK does zero-shot TTS, speech editing, enhancement, and separation from natural-language instructions.
tencent/AuK
Model deep dive · Oct 8, 2026
Breeze TTS 2 is an open-weight English and Chinese TTS model with voice cloning, text-described voices and a streaming API. Weights are non-commercial.
BreezeBlue/Breeze-TTS-2
Model deep dive · Oct 8, 2026
A contrastive scorer on frozen Qwen3-8B embeddings. It ranks the candidates you give it and answers typed questions about a state.
Contrastive-LM/CLM-v0.1-8B
Model deep dive · Oct 8, 2026
NetEase Youdao's streaming ASR model built on Qwen3-ASR: transcript text never gets rewritten, with 80 ms to 2 s chunks. Setup, code and caveats.
netease-youdao/Confucius4-R2T2
Model deep dive · Oct 8, 2026
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
Model deep dive · Oct 8, 2026
Edge0-35B-A3B streams experts from SSD so a 4-bit Qwen3.6-35B-A3B derivative decodes at about 15 tok/s in under 3 GiB of active memory.
Edge0/Edge0-35B-A3B-preview
Model deep dive · Oct 8, 2026
Google DeepMind's largest dense Gemma 4 model handles images, video frames and 256K-token context. Here's how to run it with Transformers.
google/gemma-4-31B-it
Model deep dive · Oct 8, 2026
A LoRA adapter plus a small decision head on Gemma-4-26B-A4B that gives a calibrated probability for every option in about 45 ms, and can think when it is unsure.
autotrust/GEV-26B-Decide
Model deep dive · Oct 8, 2026
A 340M classifier from Fastino that scores intent, urgency, sentiment and routing labels you choose at call time, in one forward pass, on CPU or GPU.
fastino/GLiNER2.5-Decide
Model deep dive · Oct 8, 2026
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
Model deep dive · Oct 8, 2026
A 27B vision decision model that returns a calibrated probability for every option in one forward pass. Here's how to serve it and use it.
autotrust/JEV-27B-VL
Model deep dive · Oct 8, 2026
Jina-OCR-v1 is a 570M-active-parameter OCR model that turns page images into Markdown. Here is how to run it with Transformers, vLLM, or the hosted API.
jinaai/jina-ocr-v1
Model deep dive · Oct 8, 2026
IFM's open MoE model runs 4B active parameters with a 512K context window. How to serve it, call it, and use it for agents and long documents.
IFM/K2-Horizon-MoVA-36B-A4B
Model deep dive · Oct 8, 2026
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3
Model deep dive · Oct 8, 2026
Kokoro-82M is an 82M-parameter open-weight TTS model. How to install it, generate speech, and use it for narration, pronunciation fixes, and captions.
hexgrad/Kokoro-82M
Model deep dive · Oct 8, 2026
Laya answers typed questions about text with probabilities instead of generated text. It covers 100+ languages and is Apache 2.0. Fit temperatures before trusting the probabilities.
convaiinnovations/laya
Model deep dive · Oct 8, 2026
Xiaomi's 9B SFT model, fine-tuned from Qwen3.5-9B, scores higher than its base on SWE Pro, Terminal Bench and Toolathlon. How to serve it with SGLang.
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Model deep dive · Oct 8, 2026
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
Model deep dive · Oct 8, 2026
MiniCPM5-2B is a 2.5B Llama-architecture model with 131K context and tool calling. Here's how to serve it and what it's good for.
openbmb/MiniCPM5-2B
Model deep dive · Oct 8, 2026
MiniMax H3 is an open-weight model that generates video and stereo audio together. Here's how to serve it with SGLang and where the hosted APIs fit in.
MiniMaxAI/MiniMax-H3
Model deep dive · Oct 8, 2026
NVIDIA's 100M-parameter diarizer labels up to 8 speakers, streaming with 0.32 s of buffer or offline with 30.4 s. Here's how to run it.
nvidia/Nemotron-3-Diarization
Model deep dive · Oct 8, 2026
NeoHorse-1-4B is a text-only Qwen3.5-4B fine-tune for agent harnesses, tool use and coding. Here is how to serve it with SGLang or vLLM.
TokenRhythm/NeoHorse-1-4B
Model deep dive · Oct 8, 2026
Nex-AGI's smallest Nex-N2.5 agent model, served with SGLang on 2 H100s, with tool calling and per-request thinking modes.
nex-agi/Nex-N2.5-mini
Model deep dive · Oct 8, 2026
Qwen's driving model adds a BEV perception head and a flow-matching planner to Qwen3.5-4B. Install it and sample ego trajectories from demo scenes.
Qwen/Qwen-Drive-1.0-4B
Model deep dive · Oct 8, 2026
Qwen-Image-2.1 combines text-to-image, image editing and native RGBA output in one diffusers pipeline. Here is how to run it.
Qwen/Qwen-Image-2.1
Model deep dive · Oct 8, 2026
Qwen3.8-27B is an Apache-2.0 dense 27B vision-language model with adjustable reasoning. How to serve it and call it today.
Qwen/Qwen3.8-27B
Model deep dive · Oct 8, 2026
Qwen's 125B MoE with 6B active params handles text, images and video. What it's good at, how to call it, and where it falls short.
Qwen/Qwen3.8-Flash-Next
Model deep dive · Oct 8, 2026
TeleOCR is a 1.2B Apache-2.0 model that turns text, tables, formulas and layouts into structured output, including from photos of warped pages.
XingChen-AGI/TeleOCR
Model deep dive · Oct 8, 2026
China Telecom's 29B MoE activates 4B parameters per token and has a 256K context. Here's how to call it through an OpenAI-compatible API.
XingChen-AGI/Xing4.0-29B-A4B
Model deep dive · Oct 8, 2026
YuE2-3B turns lyrics and a style prompt into a full 48 kHz stereo song. You can edit the melody and chords as an ABC score and render it again.
m-a-p/YuE2-3B
Model deep dive · Oct 7, 2026
Clef answers typed questions about any input (text, JSON, images) with a probability per option in one forward pass. Here's how to use it for routing, triage and classification.
Cloudflare/clef
Model deep dive · Oct 7, 2026
Google's new 740M multimodal embedding model runs on a CPU. Here's how to build a working semantic search over your own documents, and the prefixes that make it accurate.
google/embeddinggemma-2
Tip · Oct 7, 2026
Loading EmbeddingGemma 2 in float16 silently breaks it. Here's how to detect it and pick the right dtype for your hardware.
google/embeddinggemma-2
Model deep dive · Oct 6, 2026
Kolibri 1 is a German–English reasoning model with only 3.5B active parameters. How to serve it, control its thinking effort, and wire up tool calling.
Aleph-Alpha/Kolibri-1
Tip · Oct 5, 2026
A practical rule of thumb for truncating embeddings to cut vector-DB cost, with the numbers from EmbeddingGemma 2.