CLM-v0.1-8B: rank candidates and route agent decisions fast
A contrastive scorer on frozen Qwen3-8B embeddings. It ranks the candidates you give it and answers typed questions about a state.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
CLM-v0.1-8B is a Contrastive Language Model released by the Contrastive-LM team. It is not a chat model. It is a scorer with two small projection heads, a state head and an action head, that sit on top of a frozen Qwen3-8B encoder. The heads were trained with a bidirectional InfoNCE loss. You give it a state, such as the customer message in the card’s example, and a set of candidate actions or answers. It returns a probability distribution over those candidates. The authors call it a “System One” model. Its main selling point is speed: states and actions are encoded separately, so action embeddings can be reused. The card reports that with about 1k candidates, CLM is 13× faster than Jev, the model the card uses for its comparisons.
Key specs
| Base encoder | Qwen3-8B (frozen), last-token-pooled embeddings |
| Trainable parts | State head and action head only |
| Pre-training | ~60M Nemotron Q&A pairs |
| Mid-training | ~30M synthetic hard negatives |
| Post-training | ~1M agentic trajectories |
| Zero-shot | On par with Jev on computer-use, gaming and tool-calling tasks, up to 9× lower latency |
| Fine-tuned verifier | DeepSWE 81.6%, Terminal-Bench 2.1 87.6%, 4–6× faster than Jev |
| License | Apache 2.0 (Qwen3-8B is also Apache 2.0) |
| Language | English |
The verifier numbers come from fine-tuned heads, not from this checkpoint used zero-shot. The card says so directly, and it is covered again in the Gotchas section.
Install
You run two processes. vLLM serves Qwen3-8B as a pooling (embedding) model, and clm-serve runs the CLM API and a web playground. The card doesn’t show installing vLLM, so check that vllm is available before running the first command.
pip install contrastive-lm
# 1. encoder (Qwen3-8B embeddings)vllm serve Qwen/Qwen3-8B --served-model-name qwen3-8b --runner pooling --max-model-len 2048 --port 8090 &
# 2. API + playground at http://localhost:8700/ (fetches CLM_v0.1-8B.pt into ~/.cache/clm/)clm-serveAccording to the card, clm-serve fetches the head checkpoint (CLM_v0.1-8B.pt) into ~/.cache/clm/. Open http://localhost:8700/ to try questions in the browser.
Run it
The package has two entry points. CLMClient talks to clm-serve and answers typed questions about a state. Engine points at the vLLM embeddings endpoint and ranks free-form candidates.
from clm import CLMClient, Choice, Engine
client = CLMClient() # http://127.0.0.1:8700 by defaultr = client.system_one( state="Customer: my invoice was charged twice and nobody answers the phone!", questions={ "department": Choice(instructions="Which team should handle this?", criteria={"billing": "Charges, invoices, refunds", "technical": "Bugs and outages"}), },)print(r.answers["department"].choice)print(r.answers["department"].probabilities)
engine = Engine(emb_url="http://127.0.0.1:8090/v1/embeddings")print(engine.rank("What causes tides on Earth?", ["The Moon's gravitational pull.", "Photosynthesis in plants.", "Because the Earth is round."]))The card shows output like {'billing': 0.93878, 'technical': 0.06122} for the choice question and [{'rank': 1, 'candidate': "The Moon's gravitational pull.", 'prob': 0.993}, ...] for the ranking.
Support ticket triage
This is the card’s own example extended to a batch of tickets. The card shows three question types: Noul (spelled exactly that way in the card), Choice and Score. The card doesn’t define them. Going by its example prompts, Noul looks like a yes/no question, Choice picks one labeled option, and Score places the state on an ordered scale. You can ask several questions about the same state in one call.
from clm import CLMClient, Choice, Noul, Score
client = CLMClient()
tickets = [ "Customer: my invoice was charged twice and nobody answers the phone!", "Customer: the dashboard has shown a 500 error since this morning.", "Customer: could you send me last month's receipt when you get a chance?",]
questions = { "urgency": Noul(instructions="Is this urgent?"), "department": Choice(instructions="Which team should handle this?", criteria={"billing": "Charges, invoices, refunds", "technical": "Bugs and outages"}), "frustration": Score(instructions="How frustrated is the customer?", criteria=["Calm", "Frustrated", "Very angry"]),}
for ticket in tickets: r = client.system_one(state=ticket, questions=questions) dept = r.answers["department"] print(ticket) print(" department:", dept.choice, dept.probabilities) print(" urgency:", r.answers["urgency"]) print(" frustration:", r.answers["frustration"])The card only shows the fields of a Choice answer (.choice and .probabilities), so this script prints the Noul and Score answer objects whole instead of guessing their attribute names.
Tool routing for an agent
The card lists tool names as one kind of candidate for rank, and tool calling is one of the zero-shot tasks where it claims parity with Jev. Here the state is the user’s request and the candidates are short tool descriptions. The top-ranked tool goes to your agent loop.
from clm import Engine
engine = Engine(emb_url="http://127.0.0.1:8090/v1/embeddings")
tools = { "search_web: look up current information on the internet": "search_web", "run_sql: query the company orders database": "run_sql", "send_email: send an email to a contact": "send_email", "get_weather: current weather for a city": "get_weather",}
request = "How many orders did we ship to Germany last week?"ranked = engine.rank(request, list(tools.keys()))
# The card doesn't say the list comes back sorted, so select by the rank field.best = min(ranked, key=lambda row: row["rank"])print("chosen tool:", tools[best["candidate"]], "prob:", best["prob"])for row in sorted(ranked, key=lambda row: row["rank"]): print(row["rank"], row["prob"], row["candidate"])The probabilities only describe this set of candidates. If none of your tools fits the request, CLM will still put one of them first. Add a “no tool needed” candidate if your agent should be able to decline.
Best-of-N answer selection
The card also lists best-of-N solutions as candidates for rank. You generate N answers with any generator model, then let CLM choose one. The card makes no zero-shot accuracy claim for picking code or shell solutions, though. Its zero-shot parity claim covers computer-use, gaming and tool-calling, and its verifier results come only from fine-tuned heads. Treat the zero-shot version below as an experiment and check its picks on your own data.
from clm import Engine
engine = Engine(emb_url="http://127.0.0.1:8090/v1/embeddings")
task = "Write a shell command that counts the lines in every .py file under src/."candidates = [ "find src -name '*.py' | xargs wc -l", "ls src/*.py | wc -l", "cat src | grep py | wc", "wc -l src",]
ranked = engine.rank(task, candidates)best = min(ranked, key=lambda row: row["rank"])print("selected:", best["candidate"], best["prob"])The card’s verifier numbers on DeepSWE and Terminal-Bench 2.1 come from fine-tuned heads. Zero-shot use like the script above isn’t covered by those numbers. Only the heads are trained, so the card calls fine-tuning cheap, and it says this checkpoint is the starting point for the DeepSWE and Terminal-Bench heads. These are the card’s commands for the DeepSWE setup:
git clone https://github.com/Contrastive-LM/CLM.git && cd CLM && pip install -e .hf download Contrastive-LM/deepswe-clm-heads-8k heldout_tasks.json --local-dir heads/deepswepython train/finetune.py --task clm --init-ckpt "$(clm-download)" --out-dir runs/deepswe \ --holdout-tasks heads/deepswe/heldout_tasks.json --batch 512This block alone isn’t enough to reproduce the DeepSWE heads. It only downloads the held-out task list, and the card doesn’t show where the training trajectories come from. Get them from the fine-tuning guide linked from the card. The card also uses the hf CLI without showing how to install it, so make sure hf is available before running the download step.
Gotchas
- The encoder is fixed. The heads require Qwen3-8B last-token-pooled embeddings, so you can’t swap in a different embedding model. The card doesn’t say whether a quantized Qwen3-8B works. You also always need a running Qwen3-8B pooling server.
- It doesn’t generate text. CLM only scores the candidates you give it. If the right answer isn’t in the list, it can’t produce it.
- Probabilities are relative. The scores are normalized over the candidate set, so a 0.99 means “best of these,” not “correct.” Don’t use one fixed threshold across candidate sets of different sizes or quality.
- Zero-shot results are not verifier results. The 81.6% DeepSWE and 87.6% Terminal-Bench 2.1 figures come from fine-tuned heads, not from this checkpoint.
- Context length. The card’s vLLM command sets
--max-model-len 2048. Long states, such as full logs or large diffs, will need trimming or a different setting. The card gives no maximum input length for CLM and doesn’t say how the heads behave on longer inputs. - Caching has no documented API. The 13× speedup with ~1k candidates relies on reusing action embeddings. The card doesn’t show how to control that cache.
- English only, according to the card’s metadata.
- A newer version is announced. The card describes this model as one rung of a scaling ladder. It says a multimodal CLM-35B with stronger generalization is “coming in early October,” so it may already be out by the time you read this. Check before you commit to the 8B model.
When to pick it
Pick CLM-v0.1-8B when your agent or pipeline makes many fast decisions over a known set of options: routing tickets, choosing tools, picking a next move, or selecting among N generated answers. The Apache 2.0 license and the head-only fine-tuning make it practical to try adapting it into a verifier for your own domain. The card’s fine-tuned results cover only DeepSWE and Terminal-Bench 2.1, so you’ll need to measure how well that works on your own tasks.
Skip it if you need generated text, open-ended reasoning, or calibrated absolute confidence. Skip it if you can’t host an 8B embedding model next to your stack, because the frozen Qwen3-8B encoder is not optional. If your inputs are regularly long, note that the card’s serving command uses a 2,048-token setting and the card doesn’t document behavior beyond it, so test on your own inputs first. For a simple reranking job where a small embedding model is fast enough, this setup is more infrastructure than you need. Finally, all of the card’s speed comparisons are against Jev. If Jev isn’t what you run today, benchmark CLM against your current approach before switching.
Related
Get calibrated yes/no and multiple-choice answers with GEV-26B-Decide
A LoRA adapter plus a small decision head on Gemma-4-26B-A4B that gives a calibrated probability for every option in about 45 ms, and can think when it is unsure.
autotrust/GEV-26B-Decide
Run JEV-27B-VL for calibrated yes/no and choice decisions on images
A 27B vision decision model that returns a calibrated probability for every option in one forward pass. Here's how to serve it and use it.
autotrust/JEV-27B-VL
Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL