Run NeoHorse-1-4B as a Local Agent and Coding Model
NeoHorse-1-4B is a text-only Qwen3.5-4B fine-tune for agent harnesses, tool use and coding. Here is how to serve it with SGLang or vLLM.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
NeoHorse-1-4B is a 4B causal language model from TokenRhythm, fine-tuned from Qwen3.5-4B for text-based agent harnesses, tool use, coding and instruction following. The interesting part is how it was trained. A “routing harness” sends tasks to a mix of models, records the tool interactions and outcomes, and uses that feedback to pick the next training mixture. TokenRhythm calls this an early prototype on the way to recursive self-improvement. What you actually get is a small model under Apache-2.0. It averages 64.87 across ten benchmarks, against 58.94 for its base model. The largest single gains are on VitaBench, WorkBuddy Bench and HumanEval.
Key specs
| Property | Value |
|---|---|
| Parameters | ~4B |
| Base model | Qwen/Qwen3.5-4B (fine-tune) |
| Interface | Text in, text out. Vision weights not included |
| Context length | 262,144 native, extensible up to 1,010,000 tokens |
| Weights | Safetensors, BF16 |
| License | Apache-2.0 |
Selected results from the card (higher is better):
| Benchmark | Qwen3.5-4B | NeoHorse-1-4B | Best other model in the table |
|---|---|---|---|
| QwenClawBench | 38.47 | 44.68 | 43.52 (Spark-X2.5-4B) |
| WorkBuddy Bench | 24.62 | 34.41 | 33.37 (Agents-A1-4B) |
| PinchBench | 71.19 | 77.33 | 75.07 (Agents-A1-4B) |
| tau2-Bench | 84.29 | 88.46 | 85.08 (Nanbeige-4.2-3B) |
| BFCL v4 | 61.02 | 61.79 | 67.28 (Nanbeige-4.2-3B) |
| VitaBench | 21.50 | 32.00 | 39.25 (Agents-A1-4B) |
| HumanEval | 87.20 | 96.95 | 98.78 (Nanbeige-4.2-3B) |
| LiveCodeBench v6 | 53.71 | 59.43 | 72.50* (Nanbeige-4.2-3B) |
| IFEval | 87.06 | 88.35 | 91.13 (Spark-X2.5-4B) |
| Ten-benchmark average | 58.94 | 64.87 | 62.31 (Nanbeige-4.2-3B) |
* The card says Nanbeige’s LiveCodeBench score comes from that model’s own blog post or report, not from the same evaluation run.
NeoHorse beats its base model on all ten benchmarks. It has the top score in the table on four of them (QwenClawBench, WorkBuddy, PinchBench, tau2-Bench) and on the overall average. The gains vary a lot, from +10.50 on VitaBench down to +0.77 on BFCL v4.
Install
The card covers self-hosted serving with SGLang (the version used for the technical report) or vLLM. Both expect the checkpoint on local disk, in a directory with config.json, the tokenizer files and the weights. The card does not include a download command.
# Option A: the version used in the technical reportpip install "sglang==0.5.17"
# Option Bpip install -U vllmRun it
SGLang:
MODEL_PATH="/path/to/NeoHorse-1-4B"python3 -m sglang.launch_server \ --model-path "$MODEL_PATH" \ --served-model-name neohorse-1-4b \ --host 0.0.0.0 \ --port 30000 \ --context-length 262144 \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_codervLLM:
MODEL_PATH="/path/to/NeoHorse-1-4B"vllm serve "$MODEL_PATH" \ --served-model-name neohorse-1-4b \ --host 0.0.0.0 \ --port 8000 \ --max-model-len 262144 \ --reasoning-parser qwen3 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coderThe card says the vLLM server exposes an OpenAI-compatible /v1/chat/completions endpoint. For SGLang it shows an “OpenAI-compatible request” sent to the same path. Smoke test against vLLM:
curl http://localhost:8000/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{"model":"neohorse-1-4b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'Requests use the served model name (neohorse-1-4b), not the path on disk. For SGLang, change the port to 30000.
The Python examples below send the same request shape with requests (pip install requests). They add the sampling values from the card’s reported evaluation protocol. Two caveats apply:
top_k,min_pandrepetition_penaltyare not standard OpenAI chat-completion fields. vLLM and SGLang accept them as server-specific extra parameters. The card lists the values but never shows them in a request body. Its curl examples send onlymodel,messagesandmax_tokens.- The scripts do not set the thinking flags (
enable_thinking,force_nonempty_content), because the card does not show how to pass them per request. So the scripts do not exactly reproduce the evaluation setup.
Coding assistant
The card’s two coding benchmarks show clear gains over the base model: +9.75 on HumanEval and +5.72 on LiveCodeBench v6. This script sends a coding task and saves the reply to a file.
import sysimport requests
URL = "http://localhost:8000/v1/chat/completions"
# Values from the card's reported evaluation protocol.# top_k, min_p and repetition_penalty are vLLM/SGLang extensions, not standard OpenAI fields.SAMPLING = { "temperature": 1.0, "top_p": 0.95, "top_k": 20, "min_p": 0.0, "presence_penalty": 1.5, "repetition_penalty": 1.0,}
def ask(prompt: str, max_tokens: int = 4096) -> str: payload = { "model": "neohorse-1-4b", "messages": [{"role": "user", "content": prompt}], "max_tokens": max_tokens, **SAMPLING, } resp = requests.post(URL, json=payload, timeout=600) resp.raise_for_status() # With a reasoning parser enabled, content can be null if reasoning used up max_tokens. return resp.json()["choices"][0]["message"].get("content") or ""
if __name__ == "__main__": task = " ".join(sys.argv[1:]) or ( "Write a Python function that parses an ISO 8601 date string " "and returns the weekday name. Include three pytest tests." ) answer = ask(task) if not answer.strip(): sys.exit("Empty answer. Try a larger max_tokens.") with open("answer.md", "w") as f: f.write(answer) print(answer)Batch processing: triage a queue of tickets
This script classifies tickets by sending several requests at once. vLLM and SGLang can batch concurrent requests on the server side. That is general server behaviour, and the card says nothing about throughput or cost, so measure on your own hardware.
import jsonfrom concurrent.futures import ThreadPoolExecutorimport requests
URL = "http://localhost:8000/v1/chat/completions"LABELS = ["bug", "feature_request", "billing", "question"]
# Values from the card's reported evaluation protocol (top_k, min_p,# repetition_penalty are vLLM/SGLang extensions).SAMPLING = { "temperature": 1.0, "top_p": 0.95, "top_k": 20, "min_p": 0.0, "presence_penalty": 1.5, "repetition_penalty": 1.0,}
tickets = [ "The export button returns a 500 error since yesterday.", "Could you add dark mode to the dashboard?", "I was charged twice for my October invoice.", "How do I rotate my API key?",]
def classify(text: str) -> dict: prompt = ( f"Classify the support ticket into exactly one of: {', '.join(LABELS)}.\n" "Reply with only the label.\n\n" f"Ticket: {text}" ) payload = { "model": "neohorse-1-4b", "messages": [{"role": "user", "content": prompt}], "max_tokens": 4096, **SAMPLING, } resp = requests.post(URL, json=payload, timeout=600) resp.raise_for_status() # content can be null if reasoning used up max_tokens content = resp.json()["choices"][0]["message"].get("content") or "" raw = content.strip().strip(".`'\"").lower() label = raw if raw in LABELS else "unknown" return {"ticket": text, "label": label}
with ThreadPoolExecutor(max_workers=8) as pool: results = list(pool.map(classify, tickets))
print(json.dumps(results, indent=2))A reply only counts if, after trimming whitespace and stray punctuation, it exactly matches one of the allowed labels. Anything else becomes unknown, including an empty reply or one that mentions more than one label.
Structured extraction to JSON
IFEval and IFBench measure how well a model follows format instructions. NeoHorse improves on its base model on both (+1.29 and +5.00). This example asks for JSON, validates it, and retries once if parsing fails. If the retry also fails, it exits with an error.
import jsonimport reimport sysimport requests
URL = "http://localhost:8000/v1/chat/completions"
# Values from the card's reported evaluation protocol (top_k, min_p,# repetition_penalty are vLLM/SGLang extensions).SAMPLING = { "temperature": 1.0, "top_p": 0.95, "top_k": 20, "min_p": 0.0, "presence_penalty": 1.5, "repetition_penalty": 1.0,}
EMAIL = """Hi team, I'm Dana Ruiz from Halvorsen Logistics.We'd like 40 seats on the Pro plan starting November 1.Reach me at dana.ruiz@halvorsen.example or +1 555 0142."""
PROMPT = ( "Extract these fields from the email and return only a JSON object with keys " '"name", "company", "email", "phone", "seats" (integer), "plan", "start_date". ' "Use null for anything missing.\n\nEmail:\n" + EMAIL)
def call(messages): payload = { "model": "neohorse-1-4b", "messages": messages, "max_tokens": 4096, **SAMPLING, } resp = requests.post(URL, json=payload, timeout=600) resp.raise_for_status() # content can be null if reasoning used up max_tokens return resp.json()["choices"][0]["message"].get("content") or ""
def parse(text): match = re.search(r"\{.*\}", text, re.DOTALL) if not match: return None try: return json.loads(match.group(0)) except json.JSONDecodeError: return None
messages = [{"role": "user", "content": PROMPT}]reply = call(messages)data = parse(reply)
if data is None: messages += [ {"role": "assistant", "content": reply}, {"role": "user", "content": "That was not valid JSON. Return only the JSON object."}, ] data = parse(call(messages))
if data is None: sys.exit("No valid JSON after one retry.")
print(json.dumps(data, indent=2))Gotchas
- Text only. This release contains only the language-model weights, repackaged for text-only inference. The vision weights are not included, so do not send it images.
- Context length vs. memory. The serving commands set 262,144 tokens. The card says actual capacity depends on GPU memory and serving settings. If the server fails to start or runs out of memory, lower
--max-model-len(vLLM) or--context-length(SGLang). The card mentions extending to 1,010,000 tokens but does not explain how. - Thinking mode. The benchmarks were run with thinking enabled (
enable_thinking=true,force_nonempty_content=true), and the servers start with--reasoning-parser qwen3. With a reasoning parser, the reasoning is returned separately from the answer. If reasoning uses upmax_tokens,contentmay come back empty or null, so leave plenty of headroom and check for it. The card does not show how to toggle thinking per request. - Sampling. The reported numbers use
temperature=1.0,top_p=0.95,top_k=20,min_p=0.0,presence_penalty=1.5,repetition_penalty=1.0. Other settings may give different quality.top_k,min_pandrepetition_penaltyare server-specific request fields, not part of the standard OpenAI API. - Tool calling is enabled but not documented. The servers start with
--tool-call-parser qwen3_coder(plus--enable-auto-tool-choiceon vLLM). The card does not show a tool-call request or response, so check your server’s documentation for the exact format before you wire it into an agent. - Model name vs. path. API requests must use
--served-model-name, not the filesystem path. - Prototype status. The card calls this an initial prototype. Several agentic benchmarks (QwenClawBench, WorkBuddy Bench, PinchBench) are less well known than HumanEval or IFEval. PinchBench and VitaBench numbers come from a single run.
- License. Apache-2.0. The Alibaba Cloud copyright notice from Qwen3.5-4B is retained in the license file.
When to pick it
Pick NeoHorse-1-4B if you want a small, permissively licensed model to run inside a text-based agent harness. In the card’s table it has the top score on QwenClawBench, WorkBuddy Bench, PinchBench and tau2-Bench, and the best ten-benchmark average of the six models compared. If you already run Qwen3.5-4B for text-only work, it is worth trying. It shares the same base and scores higher on all ten reported benchmarks. It is not a drop-in replacement if you use image input, because this release has no vision weights. The card also gives no Qwen3.5-4B serving commands to compare against, so check your existing flags against the ones above.
Look elsewhere in these cases:
- Function-calling accuracy (BFCL v4): Nanbeige-4.2-3B scores 67.28, against 61.79 for NeoHorse.
- VitaBench-style agent tasks: Agents-A1-4B (39.25) and Spark-X2.5-4B (37.00) both beat NeoHorse (32.00).
- Strict instruction following: Spark-X2.5-4B leads on IFEval (91.13) and IFBench (73.33).
- Competitive coding: Nanbeige-4.2-3B reports a higher LiveCodeBench v6 score (72.50), though from its own report rather than this evaluation.
- Image input: this checkpoint cannot accept images.
- Proven agent stability: this is an early release, so test it on your own tasks before you rely on it.
Related
Run MiMo-V2.6-Distill-Qwen-9B for Coding and Agent Tasks
Xiaomi's 9B SFT model, fine-tuned from Qwen3.5-9B, scores higher than its base on SWE Pro, Terminal Bench and Toolathlon. How to serve it with SGLang.
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
Run Audio8 ASR Infinite for 24/7 Chinese and English Transcription
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
Edge0/Audio8-ASR-Infinite