CLM-v0.1-8B: rank candidates and route agent decisions fast
A contrastive scorer on frozen Qwen3-8B embeddings. It ranks the candidates you give it and answers typed questions about a state.
Contrastive-LM/CLM-v0.1-8B
A contrastive scorer on frozen Qwen3-8B embeddings. It ranks the candidates you give it and answers typed questions about a state.
Contrastive-LM/CLM-v0.1-8B
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
IFM's open MoE model runs 4B active parameters with a 512K context window. How to serve it, call it, and use it for agents and long documents.
IFM/K2-Horizon-MoVA-36B-A4B
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
Nex-AGI's smallest Nex-N2.5 agent model, served with SGLang on 2 H100s, with tool calling and per-request thinking modes.
nex-agi/Nex-N2.5-mini
Qwen3.8-27B is an Apache-2.0 dense 27B vision-language model with adjustable reasoning. How to serve it and call it today.
Qwen/Qwen3.8-27B
Qwen's 125B MoE with 6B active params handles text, images and video. What it's good at, how to call it, and where it falls short.
Qwen/Qwen3.8-Flash-Next
China Telecom's 29B MoE activates 4B parameters per token and has a 256K context. Here's how to call it through an OpenAI-compatible API.
XingChen-AGI/Xing4.0-29B-A4B