A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
A LoRA adapter plus a small decision head on Gemma-4-26B-A4B that gives a calibrated probability for every option in about 45 ms, and can think when it is unsure.
Jina-OCR-v1 is a 570M-active-parameter OCR model that turns page images into Markdown. Here is how to run it with Transformers, vLLM, or the hosted API.
Kolibri 1 is a German–English reasoning model with only 3.5B active parameters. How to serve it, control its thinking effort, and wire up tool calling.