Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Parse Documents to Markdown with Jina-OCR-v1 on One GPU

Jina-OCR-v1 is a 570M-active-parameter OCR model that turns page images into Markdown. Here is how to run it with Transformers, vLLM, or the hosted API.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Jina-OCR-v1 is an end-to-end document parsing model from Jina AI. You give it a page image and it returns clean Markdown. If you use the stricter benchmark prompt, it returns LaTeX math and HTML tables instead. It builds on DeepSeek-OCR and keeps that model’s DeepEncoder vision tower and 3B mixture-of-experts decoder, which uses only about 570M active parameters per token. Jina added post-training and a FastMTP speculative decoding head on top. What makes it notable is cost per page. Of the fourteen systems Jina measured on olmOCR-Bench, it had the highest page throughput: 2.57 pages/s on one A100 at concurrency 32. It also scores 91.14 on OmniDocBench v1.6.

Key specs

Backbone DeepSeek-OCR: DeepEncoder + 3B MoE decoder, ~570M active
Visual tokens 1024×1024 global view → 256 tokens, plus dynamic local tiles
Speculative decoding FastMTP, one dense draft head, K=3 (vLLM only)
OmniDocBench v1.6 91.14 overall
olmOCR-Bench 83.4 overall (+7.4 vs DeepSeek-OCR)
Throughput 2.57 pages/s, 1,085 output tokens/page (olmOCR-Bench, A100, concurrency 32)
License CC BY-NC 4.0

The short outputs explain much of the speed. According to the card, Surya OCR 2 produces more tokens per second (3,760 vs 2,792), but it writes 3,568 tokens per page and finishes 1.05 pages/s. Jina-OCR-v1 writes shorter transcriptions and so finishes more pages. On an NVIDIA L4, the card says FastMTP nearly doubles decoding speed compared with plain greedy decoding. Greedy verification keeps the output identical to what greedy decoding would produce.

Install

The custom modeling code ships inside the repo and loads with trust_remote_code=True. You don’t need a separate package.

Terminal window
pip install transformers torch torchvision Pillow
# for the FastMTP path:
pip install "vllm>=0.21"

The repo also includes a single example.py that covers both backends:

Terminal window
python example.py --backend transformers --image document.png
python example.py --backend vllm --image document.png

Run it

This is the minimal Transformers path, taken from the card. It processes one PIL image per call.

ocr_hf.py
import torch
from PIL import Image
from transformers import AutoModelForCausalLM, AutoProcessor
MODEL_ID = 'jinaai/jina-ocr-v1'
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID, dtype=torch.bfloat16, trust_remote_code=True,
).to(device)
image = Image.open('document.png').convert('RGB')
inputs = processor.prepare_ocr_inputs(image, device=device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
print(processor.decode_ocr(output, inputs['input_ids']))

The default prompt is:

Transcribe the provided document image into a clean Markdown format, preserving the natural reading order.

Transcribing a folder of scans with vLLM

vLLM is the only backend that runs the FastMTP draft head. Call register() once per process, before you construct LLM. Separately, put the snapshot directory on sys.path so worker processes can import deepseek_ocr_mtp. The script below loops over a directory of page images and writes one .md file per page.

It is a sequential loop: each llm.chat call handles one page and waits for it to finish, so pages are not batched together. Don’t expect it to reach the card’s 2.57 pages/s, which was measured at concurrency 32. The card doesn’t show how to submit many pages at once.

folder_vllm.py
import sys
from pathlib import Path
from huggingface_hub import snapshot_download
from PIL import Image
from vllm import LLM
sys.path.insert(0, snapshot_download('jinaai/jina-ocr-v1'))
from deepseek_ocr_mtp import DEFAULT_OCR_PROMPT, log_spec_stats, register, vllm_llm_kwargs, vllm_sampling_params
register()
llm = LLM(
**vllm_llm_kwargs(
'jinaai/jina-ocr-v1',
num_speculative_tokens=3, # K; 0 disables MTP
mtp_heads=1,
mtp_recursive=True,
)
)
params = vllm_sampling_params(max_tokens=4096)
in_dir = Path('scans')
out_dir = Path('markdown')
out_dir.mkdir(exist_ok=True)
for path in sorted(in_dir.glob('*.png')):
image = Image.open(path).convert('RGB')
outputs = llm.chat(
[{
'role': 'user',
'content': [
{'type': 'image_pil', 'image_pil': image},
{'type': 'text', 'text': DEFAULT_OCR_PROMPT},
],
}],
sampling_params=params,
)
(out_dir / f'{path.stem}.md').write_text(outputs[0].outputs[0].text)
print(f'done: {path.name}')
log_spec_stats(llm)

vllm_sampling_params() sets temperature=0.0, repetition_penalty=1.05, and vLLM’s built-in n-gram repetition stop (max_pattern_size=35, min_pattern_size=35, min_count=10). That repetition stop works together with FastMTP. log_spec_stats(llm) flushes vLLM’s SpecDecoding metrics after a short run. Without it, vLLM prints them on a roughly 10-second interval.

Extracting math and tables with the benchmark prompt

The benchmark prompt is stricter than the default. It asks for LaTeX math with $/$$ delimiters and HTML tables with <th>, rowspan, and colspan, and it drops headers, footers, and figures. Use it if your downstream pipeline needs math and tables in that structured form. Pass it with the prompt= argument of prepare_ocr_inputs.

strict_extract.py
import torch
from PIL import Image
from transformers import AutoModelForCausalLM, AutoProcessor
MODEL_ID = 'jinaai/jina-ocr-v1'
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
STRICT_PROMPT = """Just return the plain text representation of this document as if you were reading it naturally.
Turn equations and math symbols into a LaTeX representation, make sure to use $ and $ as a delimiter for inline math, and $$ and $$ for block math. Do NOT use ascii or unicode math symbols such as ∈ ∉ ⊂ ⊃ ⊆ ⊇ ∅ ∪ ∩ ∀ ∃ ¬, just use LaTeX syntax, ex $ \\in $ $ \\notin $ etc. If you were going to surround a math expression in \\( \\) or \\[ \\] delimiters, surround it with $ $ or $$ $$ instead.
Convert tables into HTML format. Keep the syntax simple, but use <th> for header rows, and use rowspan and colspans appropriately. Don't use <br> inside of table cells, just split that into new rows as needed. Do NOT use LaTeX or Markdown table syntax.
Ignore all graphical content in the image document. Do not attempt to describe or convert the images.
Remove the headers and footers, but keep references and footnotes.
Read any natural handwriting.
This is likely one page out of several in the document, so be sure to preserve any sentences that come from the previous page, or continue onto the next page, exactly as they are.
If there is no text at all that you think you should read, you can output null."""
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID, dtype=torch.bfloat16, trust_remote_code=True,
).to(device)
image = Image.open('paper_page.png').convert('RGB')
inputs = processor.prepare_ocr_inputs(image, prompt=STRICT_PROMPT, device=device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
text = processor.decode_ocr(output, inputs['input_ids'])
with open('paper_page.txt', 'w') as f:
f.write(text)
print(text)

The prompt tells the model to output null for pages with no readable text. Filter those out before you embed anything.

Multi-page PDFs without a GPU

If you can’t host the model yourself, Jina Reader fetches a PDF by URL, renders it, runs jina-ocr-v1, and returns Markdown. Set X-Page (1-indexed) to request one page at a time. The card doesn’t say what Reader returns for an X-Page past the end of the document, so the script takes the page count as an argument rather than guessing. It stops on any non-200 response and does not save error bodies as .md files.

pdf_pages.sh
#!/usr/bin/env bash
set -euo pipefail
PDF_URL="https://example.com/document.pdf"
PAGES="${1:?usage: $0 <page-count>}"
for page in $(seq 1 "$PAGES"); do
tmp="page_${page}.tmp"
code=$(curl -s -o "$tmp" -w "%{http_code}" \
"https://r.jina.ai/${PDF_URL}" \
-H "Authorization: Bearer $JINA_API_KEY" \
-H "X-Respond-With: jina-ocr-v1" \
-H "X-Page: ${page}")
if [ "$code" != "200" ]; then
echo "page ${page}: HTTP ${code}, stopping" >&2
cat "$tmp" >&2
rm -f "$tmp"
exit 1
fi
mv "$tmp" "page_${page}.md"
echo "page ${page}: done"
done

For a single image, there is also an OpenAI-compatible endpoint at https://api.jina.ai/v1/chat/completions with "model": "jina-ocr-v1". It accepts an image URL or a data:image/...;base64,... string, and you can add "stream": true to stream the response. The card says this endpoint returns HTTP 503 on a cold start and that you should retry after 30–60 seconds.

Gotchas

  • Non-commercial license. CC BY-NC 4.0 rules out commercial use unless you contact Jina sales. Check this before anything else.
  • No speedup in Transformers. generate() there runs only the MoE decoder. The MTP weights (mtp_module.*, mtp_embed_tokens.*) are ignored on load. To get the speedup you need vLLM ≥ 0.21.
  • vLLM must use method="eagle". vllm_llm_kwargs() sets this for you. If you build the speculative config by hand with vLLM’s default method="mtp", the recursive draft head FastMTP was trained with no longer works correctly.
  • Two setup steps in vLLM. Call register() once per process before LLM(...). Separately, put the snapshot directory on sys.path, or the worker processes can’t import deepseek_ocr_mtp.
  • Don’t override the generation config. eos/pad come from generation_config.json, so don’t pass a fresh GenerationConfig. Also don’t pass use_fast=False, because the tokenizer is LlamaTokenizerFast.
  • Repetition guard. In Transformers, generate() attaches a sliding-window no-repeat-n-gram processor (no_repeat_ngram_size=35, ngram_window=1024) and whitelists the <td>/</td> tokens. You can override those keyword arguments, or set no_repeat_ngram_size=0 to disable it.
  • One image per Transformers call. vLLM is the backend with FastMTP.
  • Benchmark prompt ≠ default prompt. The 91.14 and 83.4 scores were measured with the strict prompt. The default Markdown prompt is a different setting, and the card reports no scores for it.

When to pick it

Pick Jina-OCR-v1 for high-volume document parsing where cost per page matters: research papers, reports, and scanned archives going into a search or RAG pipeline. In Jina’s olmOCR-Bench measurement it ran at 2.57 pages/s on one A100, and the card says FastMTP nearly doubles decoding speed on an L4. The card doesn’t list memory requirements or minimum hardware. Only about 570M parameters are active per token, but the full 3B MoE still has to be loaded, so check that it fits your GPU. If your downstream code needs structure, the benchmark prompt gives you LaTeX math and HTML tables.

Skip it if you are building a commercial product and can’t get a license from Jina. Skip it if you need figure or chart descriptions, because the benchmark prompt explicitly ignores graphical content and the card doesn’t evaluate image understanding. The card only documents OCR prompts, so don’t plan on using it as a general vision-language model. And if you only have Transformers available, you lose the speculative decoding advantage, so compare it against other OCR models on your own pages before you commit.

Related