Parse Documents to Markdown with Jina-OCR-v1 on One GPU
Jina-OCR-v1 is a 570M-active-parameter OCR model that turns page images into Markdown. Here is how to run it with Transformers, vLLM, or the hosted API.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Jina-OCR-v1 is an end-to-end document parsing model from Jina AI. You give it a page image and it returns clean Markdown. If you use the stricter benchmark prompt, it returns LaTeX math and HTML tables instead. It builds on DeepSeek-OCR and keeps that model’s DeepEncoder vision tower and 3B mixture-of-experts decoder, which uses only about 570M active parameters per token. Jina added post-training and a FastMTP speculative decoding head on top. What makes it notable is cost per page. Of the fourteen systems Jina measured on olmOCR-Bench, it had the highest page throughput: 2.57 pages/s on one A100 at concurrency 32. It also scores 91.14 on OmniDocBench v1.6.
Key specs
| Backbone | DeepSeek-OCR: DeepEncoder + 3B MoE decoder, ~570M active |
| Visual tokens | 1024×1024 global view → 256 tokens, plus dynamic local tiles |
| Speculative decoding | FastMTP, one dense draft head, K=3 (vLLM only) |
| OmniDocBench v1.6 | 91.14 overall |
| olmOCR-Bench | 83.4 overall (+7.4 vs DeepSeek-OCR) |
| Throughput | 2.57 pages/s, 1,085 output tokens/page (olmOCR-Bench, A100, concurrency 32) |
| License | CC BY-NC 4.0 |
The short outputs explain much of the speed. According to the card, Surya OCR 2 produces more tokens per second (3,760 vs 2,792), but it writes 3,568 tokens per page and finishes 1.05 pages/s. Jina-OCR-v1 writes shorter transcriptions and so finishes more pages. On an NVIDIA L4, the card says FastMTP nearly doubles decoding speed compared with plain greedy decoding. Greedy verification keeps the output identical to what greedy decoding would produce.
Install
The custom modeling code ships inside the repo and loads with trust_remote_code=True. You don’t need a separate package.
pip install transformers torch torchvision Pillow# for the FastMTP path:pip install "vllm>=0.21"The repo also includes a single example.py that covers both backends:
python example.py --backend transformers --image document.pngpython example.py --backend vllm --image document.pngRun it
This is the minimal Transformers path, taken from the card. It processes one PIL image per call.
import torchfrom PIL import Imagefrom transformers import AutoModelForCausalLM, AutoProcessor
MODEL_ID = 'jinaai/jina-ocr-v1'device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)model = AutoModelForCausalLM.from_pretrained( MODEL_ID, dtype=torch.bfloat16, trust_remote_code=True,).to(device)
image = Image.open('document.png').convert('RGB')inputs = processor.prepare_ocr_inputs(image, device=device)output = model.generate(**inputs, max_new_tokens=4096, do_sample=False)print(processor.decode_ocr(output, inputs['input_ids']))The default prompt is:
Transcribe the provided document image into a clean Markdown format, preserving the natural reading order.Transcribing a folder of scans with vLLM
vLLM is the only backend that runs the FastMTP draft head. Call register() once per process, before you construct LLM. Separately, put the snapshot directory on sys.path so worker processes can import deepseek_ocr_mtp. The script below loops over a directory of page images and writes one .md file per page.
It is a sequential loop: each llm.chat call handles one page and waits for it to finish, so pages are not batched together. Don’t expect it to reach the card’s 2.57 pages/s, which was measured at concurrency 32. The card doesn’t show how to submit many pages at once.
import sysfrom pathlib import Path
from huggingface_hub import snapshot_downloadfrom PIL import Imagefrom vllm import LLM
sys.path.insert(0, snapshot_download('jinaai/jina-ocr-v1'))
from deepseek_ocr_mtp import DEFAULT_OCR_PROMPT, log_spec_stats, register, vllm_llm_kwargs, vllm_sampling_params
register()
llm = LLM( **vllm_llm_kwargs( 'jinaai/jina-ocr-v1', num_speculative_tokens=3, # K; 0 disables MTP mtp_heads=1, mtp_recursive=True, ))params = vllm_sampling_params(max_tokens=4096)
in_dir = Path('scans')out_dir = Path('markdown')out_dir.mkdir(exist_ok=True)
for path in sorted(in_dir.glob('*.png')): image = Image.open(path).convert('RGB') outputs = llm.chat( [{ 'role': 'user', 'content': [ {'type': 'image_pil', 'image_pil': image}, {'type': 'text', 'text': DEFAULT_OCR_PROMPT}, ], }], sampling_params=params, ) (out_dir / f'{path.stem}.md').write_text(outputs[0].outputs[0].text) print(f'done: {path.name}')
log_spec_stats(llm)vllm_sampling_params() sets temperature=0.0, repetition_penalty=1.05, and vLLM’s built-in n-gram repetition stop (max_pattern_size=35, min_pattern_size=35, min_count=10). That repetition stop works together with FastMTP. log_spec_stats(llm) flushes vLLM’s SpecDecoding metrics after a short run. Without it, vLLM prints them on a roughly 10-second interval.
Extracting math and tables with the benchmark prompt
The benchmark prompt is stricter than the default. It asks for LaTeX math with $/$$ delimiters and HTML tables with <th>, rowspan, and colspan, and it drops headers, footers, and figures. Use it if your downstream pipeline needs math and tables in that structured form. Pass it with the prompt= argument of prepare_ocr_inputs.
import torchfrom PIL import Imagefrom transformers import AutoModelForCausalLM, AutoProcessor
MODEL_ID = 'jinaai/jina-ocr-v1'device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
STRICT_PROMPT = """Just return the plain text representation of this document as if you were reading it naturally.Turn equations and math symbols into a LaTeX representation, make sure to use $ and $ as a delimiter for inline math, and $$ and $$ for block math. Do NOT use ascii or unicode math symbols such as ∈ ∉ ⊂ ⊃ ⊆ ⊇ ∅ ∪ ∩ ∀ ∃ ¬, just use LaTeX syntax, ex $ \\in $ $ \\notin $ etc. If you were going to surround a math expression in \\( \\) or \\[ \\] delimiters, surround it with $ $ or $$ $$ instead.Convert tables into HTML format. Keep the syntax simple, but use <th> for header rows, and use rowspan and colspans appropriately. Don't use <br> inside of table cells, just split that into new rows as needed. Do NOT use LaTeX or Markdown table syntax.Ignore all graphical content in the image document. Do not attempt to describe or convert the images.Remove the headers and footers, but keep references and footnotes.Read any natural handwriting.This is likely one page out of several in the document, so be sure to preserve any sentences that come from the previous page, or continue onto the next page, exactly as they are.If there is no text at all that you think you should read, you can output null."""
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)model = AutoModelForCausalLM.from_pretrained( MODEL_ID, dtype=torch.bfloat16, trust_remote_code=True,).to(device)
image = Image.open('paper_page.png').convert('RGB')inputs = processor.prepare_ocr_inputs(image, prompt=STRICT_PROMPT, device=device)output = model.generate(**inputs, max_new_tokens=4096, do_sample=False)text = processor.decode_ocr(output, inputs['input_ids'])
with open('paper_page.txt', 'w') as f: f.write(text)print(text)The prompt tells the model to output null for pages with no readable text. Filter those out before you embed anything.
Multi-page PDFs without a GPU
If you can’t host the model yourself, Jina Reader fetches a PDF by URL, renders it, runs jina-ocr-v1, and returns Markdown. Set X-Page (1-indexed) to request one page at a time. The card doesn’t say what Reader returns for an X-Page past the end of the document, so the script takes the page count as an argument rather than guessing. It stops on any non-200 response and does not save error bodies as .md files.
#!/usr/bin/env bashset -euo pipefail
PDF_URL="https://example.com/document.pdf"PAGES="${1:?usage: $0 <page-count>}"
for page in $(seq 1 "$PAGES"); do tmp="page_${page}.tmp" code=$(curl -s -o "$tmp" -w "%{http_code}" \ "https://r.jina.ai/${PDF_URL}" \ -H "Authorization: Bearer $JINA_API_KEY" \ -H "X-Respond-With: jina-ocr-v1" \ -H "X-Page: ${page}") if [ "$code" != "200" ]; then echo "page ${page}: HTTP ${code}, stopping" >&2 cat "$tmp" >&2 rm -f "$tmp" exit 1 fi mv "$tmp" "page_${page}.md" echo "page ${page}: done"doneFor a single image, there is also an OpenAI-compatible endpoint at https://api.jina.ai/v1/chat/completions with "model": "jina-ocr-v1". It accepts an image URL or a data:image/...;base64,... string, and you can add "stream": true to stream the response. The card says this endpoint returns HTTP 503 on a cold start and that you should retry after 30–60 seconds.
Gotchas
- Non-commercial license. CC BY-NC 4.0 rules out commercial use unless you contact Jina sales. Check this before anything else.
- No speedup in Transformers.
generate()there runs only the MoE decoder. The MTP weights (mtp_module.*,mtp_embed_tokens.*) are ignored on load. To get the speedup you need vLLM ≥ 0.21. - vLLM must use
method="eagle".vllm_llm_kwargs()sets this for you. If you build the speculative config by hand with vLLM’s defaultmethod="mtp", the recursive draft head FastMTP was trained with no longer works correctly. - Two setup steps in vLLM. Call
register()once per process beforeLLM(...). Separately, put the snapshot directory onsys.path, or the worker processes can’t importdeepseek_ocr_mtp. - Don’t override the generation config.
eos/padcome fromgeneration_config.json, so don’t pass a freshGenerationConfig. Also don’t passuse_fast=False, because the tokenizer isLlamaTokenizerFast. - Repetition guard. In Transformers,
generate()attaches a sliding-window no-repeat-n-gram processor (no_repeat_ngram_size=35,ngram_window=1024) and whitelists the<td>/</td>tokens. You can override those keyword arguments, or setno_repeat_ngram_size=0to disable it. - One image per Transformers call. vLLM is the backend with FastMTP.
- Benchmark prompt ≠ default prompt. The 91.14 and 83.4 scores were measured with the strict prompt. The default Markdown prompt is a different setting, and the card reports no scores for it.
When to pick it
Pick Jina-OCR-v1 for high-volume document parsing where cost per page matters: research papers, reports, and scanned archives going into a search or RAG pipeline. In Jina’s olmOCR-Bench measurement it ran at 2.57 pages/s on one A100, and the card says FastMTP nearly doubles decoding speed on an L4. The card doesn’t list memory requirements or minimum hardware. Only about 570M parameters are active per token, but the full 3B MoE still has to be loaded, so check that it fits your GPU. If your downstream code needs structure, the benchmark prompt gives you LaTeX math and HTML tables.
Skip it if you are building a commercial product and can’t get a license from Jina. Skip it if you need figure or chart descriptions, because the benchmark prompt explicitly ignores graphical content and the card doesn’t evaluate image understanding. The card only documents OCR prompts, so don’t plan on using it as a general vision-language model. And if you only have Transformers available, you lose the speculative decoding advantage, so compare it against other OCR models on your own pages before you commit.
Related
Parse Scanned and Phone-Photo Documents with TeleOCR (1.2B)
TeleOCR is a 1.2B Apache-2.0 model that turns text, tables, formulas and layouts into structured output, including from photos of warped pages.
XingChen-AGI/TeleOCR
Qwen3.8-Flash-Next: Serve a 6B-Active Multimodal Agent Model
Qwen's 125B MoE with 6B active params handles text, images and video. What it's good at, how to call it, and where it falls short.
Qwen/Qwen3.8-Flash-Next
Run Audio8 ASR Infinite for 24/7 Chinese and English Transcription
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
Edge0/Audio8-ASR-Infinite