Parse Scanned and Phone-Photo Documents with TeleOCR (1.2B)
TeleOCR is a 1.2B Apache-2.0 model that turns text, tables, formulas and layouts into structured output, including from photos of warped pages.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
TeleOCR is a ~1.2B-parameter vision-language model for document parsing, released under Apache-2.0. Its authors are Peng Cai and colleagues. The model was first published as NaviDC-OCR under StarDoc-AI. Most document parsers handle either clean digital pages or camera-captured pages. TeleOCR handles both with one model. According to the card, it parses layout and content on distorted documents without a separate dewarping or rectification step. The card’s evidence for this is visual examples on the DocUNet and DIR300 datasets, with no scores. One set of weights covers plain text, tables, LaTeX formulas, code, page layout and extracting the table implied by a scientific figure. You choose the task with a fixed prompt string.
Key specs
- Size: ~1.2B parameters. The card doesn’t describe the architecture or backbone. Its acknowledgements list MinerU, Qwen2.5-VL, Qwen3, Transformers, PyTorch and FlashAttention.
- Languages: the card’s metadata lists Chinese and English
- Outputs: plain text, tables in OTSL format (the card includes an HTML converter), LaTeX formulas, code, layout analysis, and tables extracted from scientific figures
- License: Apache-2.0
The authors report these benchmark numbers on the card. None of them come from independent evaluation:
| Benchmark | TeleOCR | Closest competitor on the card |
|---|---|---|
| OmniDocBench v1.6 overall | 96.87 | OvisOCR2 (0.8B): 96.58 |
| OmniDocBench v1.6 Table TEDS | 97.05 | OvisOCR2 / PaddleOCR-VL-1.6 / HunyuanOCR-1.5: 94.76 |
| Wild_OmniDocBench overall | 88.53 | OvisOCR2: 87.91 |
| PureDocBench real-degraded overall | 70.85 | Gemini-3.1-Pro: 71.98 (ahead of TeleOCR) |
| Dr.DocBench Challenge overall | 67.96 | MinerU 2.5 Pro: 62.26 |
In the card’s OmniDocBench v1.6 table, the HunyuanOCR-1.5 row matches the PaddleOCR-VL-1.6 row in every column except Overall (94.74 vs 96.33). The other five metrics are identical (0.033 / 97.49 / 94.76 / 97.11 / 0.127). That may be a copy error, so treat comparisons against either model with care.
Results differ by benchmark:
- Tables. TeleOCR has the top Table TEDS on OmniDocBench v1.6 (97.05) and Wild_OmniDocBench (89.05). Other benchmarks show something different. On Dr.DocBench its Table TEDS is 64.97, behind MinerU 2.5 Pro at 67.75. On PureDocBench’s digital-degraded table score it gets 80.45, behind OvisOCR2 (84.71), FD-RL (83.22) and MinerU2.5-Pro (80.73).
- Formulas. On OmniDocBench v1.6 its Formula CDM is 96.36. That is below OvisOCR2 (97.53), PaddleOCR-VL-1.6 and HunyuanOCR-1.5 (97.49), MinerU2.5-Pro (97.45), GLM-OCR (97.18) and PaddleOCR-VL-1.5 (96.69). On Wild_OmniDocBench, five listed models score higher than its 88.26. PureDocBench is different: TeleOCR has the top formula score on the clean and digital-degraded splits. On the real-degraded split, Gemini-3.1-Pro (68.62) beats it (65.11).
- Real-degraded photos. On PureDocBench’s real-degraded split, Gemini-3.1-Pro has the higher overall score.
Install
pip install transformers torch pillowRun it
The card loads the model through AutoModel with trust_remote_code=True and wraps generation in an infer() helper. The examples below import this file. It loads the model at import time, so each script that imports it loads the model once when it starts.
import torchfrom PIL import Imagefrom transformers import AutoProcessor, AutoModel
# The id used in the card's own code. The card itself is hosted at XingChen-AGI/TeleOCR.MODEL_ID = "StarDoc-AI/TeleOCR"
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True, use_fast=True)model = AutoModel.from_pretrained( MODEL_ID, trust_remote_code=True, torch_dtype=torch.bfloat16,).cuda().eval()
PROMPT_TEXT = "Please output the text content from the image."PROMPT_TABLE = "This is the image of a table. Please output the table in OTSL format."PROMPT_FORMULA = "Please write out the expression of the formula in the image using LaTeX format."PROMPT_CODE = "The image contains a code snippet, please output the parsing result."PROMPT_LAYOUT = "Analyze the image layout."PROMPT_LAYOUT_DISTORTED = "\nMulti-point Layout Segmentation Analysis."PROMPT_FIGURE = "This is a scientific figure. Please extract the table implied by this figure."
def infer(image: Image.Image, prompt: str) -> str: messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": [ {"type": "image"}, {"type": "text", "text": prompt}, ]}, ] chat_prompt = processor.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, ) inputs = processor( text=[chat_prompt], images=[image.convert("RGB")], padding=True, return_tensors="pt", ).to(device=model.device, dtype=model.dtype) output_ids = model.generate( **inputs, use_cache=True, max_new_tokens=4096, do_sample=False, ) output_ids = output_ids.cpu().tolist()[0][len(inputs.input_ids[0]):] return processor.batch_decode( [output_ids], skip_special_tokens=True, clean_up_tokenization_spaces=False, )[0].strip()
if __name__ == "__main__": import sys
image = Image.open(sys.argv[1]).convert("RGB") print(infer(image, PROMPT_TEXT))python teleocr.py page.pngRun the text prompt over a folder of images
This script runs the text prompt on every image in a folder and saves one .txt file per image. It processes one image at a time, as the card does. The card only demonstrates the text prompt on a single sample image. It doesn’t show results on full pages, and it points to the GitHub repo for complete document parsing. Check the output on your own pages before you rely on it.
import sysfrom pathlib import Path
from PIL import Image
from teleocr import infer, PROMPT_TEXT
src = Path(sys.argv[1])dst = Path(sys.argv[2])dst.mkdir(parents=True, exist_ok=True)
for path in sorted(src.iterdir()): if path.suffix.lower() not in {".png", ".jpg", ".jpeg"}: continue image = Image.open(path).convert("RGB") text = infer(image, PROMPT_TEXT) (dst / f"{path.stem}.txt").write_text(text, encoding="utf-8") print(f"{path.name}: {len(text)} chars")python batch_text.py ./scans ./outExtract tables to HTML
TeleOCR scores highest on tables in OmniDocBench v1.6 (97.05 TEDS) and Wild_OmniDocBench (89.05 TEDS). As shown above, other benchmarks on the card rank it lower. The model outputs tables in OTSL, a compact token format, not HTML. The OTSL-to-HTML converter below is copied from the card. The card also runs it on the output of the scientific-figure prompt, which asks for the table implied by a figure.
import htmlimport itertoolsimport refrom dataclasses import dataclass
@dataclassclass TableCell: text: str start_row_offset_idx: int end_row_offset_idx: int start_col_offset_idx: int end_col_offset_idx: int row_span: int = 1 col_span: int = 1
OTSL_NL = "<nl>"OTSL_FCEL = "<fcel>"OTSL_ECEL = "<ecel>"OTSL_LCEL = "<lcel>"OTSL_UCEL = "<ucel>"OTSL_XCEL = "<xcel>"OTSL_TOKENS = [OTSL_NL, OTSL_FCEL, OTSL_ECEL, OTSL_LCEL, OTSL_UCEL, OTSL_XCEL]
def _otsl_extract_tokens_and_text(text: str): pattern = "(" + "|".join(map(re.escape, OTSL_TOKENS)) + ")" tokens = re.findall(pattern, text) parts = [part for part in re.split(pattern, text) if part.strip()] return tokens, parts
def _count_right(rows, row_idx, col_idx, tokens): span = 0 while col_idx < len(rows[row_idx]) and rows[row_idx][col_idx] in tokens: span += 1 col_idx += 1 return span
def _count_down(rows, row_idx, col_idx, tokens): span = 0 while row_idx < len(rows) and col_idx < len(rows[row_idx]) and rows[row_idx][col_idx] in tokens: span += 1 row_idx += 1 return span
def _otsl_parse_texts(parts, tokens): rows = [list(row) for is_nl, row in itertools.groupby(tokens, lambda token: token == OTSL_NL) if not is_nl] if not rows: return [], []
max_cols = max(len(row) for row in rows) for row in rows: row.extend([OTSL_ECEL] * (max_cols - len(row)))
cells = [] row_idx = 0 col_idx = 0 for idx, part in enumerate(parts): if part in (OTSL_FCEL, OTSL_ECEL): cell_text = "" right_offset = 1 if part != OTSL_ECEL and idx + 1 < len(parts) and parts[idx + 1] not in OTSL_TOKENS: cell_text = parts[idx + 1].strip() right_offset = 2
next_right = parts[idx + right_offset] if idx + right_offset < len(parts) else "" next_bottom = rows[row_idx + 1][col_idx] if row_idx + 1 < len(rows) and col_idx < len(rows[row_idx + 1]) else "" col_span = 1 + (_count_right(rows, row_idx, col_idx + 1, {OTSL_LCEL, OTSL_XCEL}) if next_right in {OTSL_LCEL, OTSL_XCEL} else 0) row_span = 1 + (_count_down(rows, row_idx + 1, col_idx, {OTSL_UCEL, OTSL_XCEL}) if next_bottom in {OTSL_UCEL, OTSL_XCEL} else 0) cells.append(TableCell( text=cell_text, row_span=row_span, col_span=col_span, start_row_offset_idx=row_idx, end_row_offset_idx=row_idx + row_span, start_col_offset_idx=col_idx, end_col_offset_idx=col_idx + col_span, )) if part in (OTSL_FCEL, OTSL_ECEL, OTSL_LCEL, OTSL_UCEL, OTSL_XCEL): col_idx += 1 elif part == OTSL_NL: row_idx += 1 col_idx = 0 return cells, rows
def convert_otsl_to_html(otsl_content: str) -> str: if otsl_content.startswith("<table") and otsl_content.endswith("</table>"): return otsl_content
tokens, parts = _otsl_extract_tokens_and_text(otsl_content) cells, rows = _otsl_parse_texts(parts, tokens) if not cells or not rows: return ""
grid = [[None for _ in range(len(rows[0]))] for _ in range(len(rows))] for cell in cells: for row_idx in range(cell.start_row_offset_idx, min(cell.end_row_offset_idx, len(rows))): for col_idx in range(cell.start_col_offset_idx, min(cell.end_col_offset_idx, len(rows[0]))): grid[row_idx][col_idx] = cell
html_rows = [] for row_idx, row in enumerate(grid): html_rows.append("<tr>") for col_idx, cell in enumerate(row): if cell is None or cell.start_row_offset_idx != row_idx or cell.start_col_offset_idx != col_idx: continue attrs = "" if cell.row_span > 1: attrs += f' rowspan="{cell.row_span}"' if cell.col_span > 1: attrs += f' colspan="{cell.col_span}"' html_rows.append(f"<td{attrs}>{html.escape(cell.text.strip())}</td>") html_rows.append("</tr>") return "<table>" + "".join(html_rows) + "</table>"import sys
from PIL import Image
from otsl import convert_otsl_to_htmlfrom teleocr import infer, PROMPT_TABLE, PROMPT_FIGURE
PROMPTS = {"table": PROMPT_TABLE, "figure": PROMPT_FIGURE}
mode, path = sys.argv[1], sys.argv[2]if mode not in PROMPTS: sys.exit(f"Unknown mode {mode!r}: use 'table' or 'figure'")prompt = PROMPTS[mode]
image = Image.open(path).convert("RGB")raw_otsl = infer(image, prompt)table_html = convert_otsl_to_html(raw_otsl)
if not table_html: print("Could not parse OTSL output:\n" + raw_otsl)else: print(table_html)python tables.py table invoice_table.png > table.htmlpython tables.py figure scientific_figure.png > figure_data.htmlFormulas to LaTeX
The formula prompt returns LaTeX. The card passes that output through a post_process step that strips \[ ... \] and wraps the result in $$...$$. The ContentBlock dataclass and post_process function below are copied from the card. The formula call follows the card’s example, which wraps the whole image in one equation block. Keep in mind that on the card’s OmniDocBench v1.6 and Wild_OmniDocBench results, several models score higher than TeleOCR on formulas.
import sysfrom dataclasses import dataclass
from PIL import Image
from otsl import convert_otsl_to_htmlfrom teleocr import infer, PROMPT_FORMULA
@dataclassclass ContentBlock: type: str bbox: list[float] angle: int | None = None content: str | None = None
def post_process(blocks: list[ContentBlock]) -> list[ContentBlock]: for block in blocks: if block.type == "table" and block.content: block.content = convert_otsl_to_html(block.content) elif block.type == "equation" and block.content: content = block.content.strip() content = content.removeprefix("\\[").removesuffix("\\]").strip() if not (content.startswith("$") and content.endswith("$")): content = f"$${content}$$" block.content = content return [block for block in blocks if block.type != "equation_block"]
if __name__ == "__main__": image = Image.open(sys.argv[1]).convert("RGB") raw_formula = infer(image, PROMPT_FORMULA) formula_block = ContentBlock("equation", [0.0, 0.0, 1.0, 1.0], content=raw_formula) print(post_process([formula_block])[0].content)python formula.py formula.pngLayout analysis on distorted pages
The card has two layout prompts: one for regular pages and one for distorted documents. The distorted-document prompt starts with a newline. Before running either one, the card resizes the image to 1036×1036 with bicubic resampling. The card supports its distorted-document claims with visual examples on DocUNet and DIR300, not scores. The card also doesn’t document the layout output format, so this example prints the raw result. Check it on your own data before you write a parser for it.
import sys
from PIL import Image
from teleocr import infer, PROMPT_LAYOUT, PROMPT_LAYOUT_DISTORTED
path = sys.argv[1]distorted = len(sys.argv) > 2 and sys.argv[2] == "--distorted"
image = Image.open(path).convert("RGB")image = image.resize((1036, 1036), Image.Resampling.BICUBIC)
prompt = PROMPT_LAYOUT_DISTORTED if distorted else PROMPT_LAYOUTprint(infer(image, prompt).strip())python layout.py scanned_page.jpgpython layout.py phone_photo.jpg --distortedGotchas
- Repo id mismatch. The card is hosted at
XingChen-AGI/TeleOCR, but its code loadsStarDoc-AI/TeleOCR, and that is the id these examples use. The model was also renamed from NaviDC-OCR. We haven’t checked which repo id loads withtrust_remote_code. If one doesn’t work, try the other. trust_remote_code=Trueis required. The model runs custom code from the repo, so review that code before you deploy it.- The example assumes a CUDA GPU in bf16. It calls
.cuda()and casts inputs tomodel.dtype. The card doesn’t document any other hardware path. A community GGUF conversion with llama.cpp support exists. It was made from NaviDC-OCR before the rename, and the card doesn’t say whether it matches the current TeleOCR weights or accepts the same prompts. - Use the exact prompts. Tasks are selected by the fixed prompt strings shown above. The distorted-layout prompt starts with
\n, so keep it. - Resize images for layout. The card resizes to 1036×1036 for both layout prompts.
- Tables come out as OTSL, not HTML. Run the output through
convert_otsl_to_html. Formulas need the card’spost_processstep (included informula.pyabove) if you want$$...$$delimiters. max_new_tokens=4096. Very dense inputs may get cut off at this limit.- Model loads on import.
teleocr.pyloads the weights when it’s imported. Each script above loads the model once per run. - Single-region tasks. The card shows one task per image. For complete document parsing, the card points to the GitHub repo and doesn’t describe the pipeline itself.
- Languages. The card’s metadata lists only Chinese (
zh) and English (en). It doesn’t say anything about other languages, so test them yourself if you need them. - Self-reported benchmarks. Every score in this post comes from the authors’ card.
When to pick it
Pick TeleOCR if you need a small, Apache-2.0 parser for Chinese or English documents, especially if some input arrives as camera photos instead of clean PDFs. On the card it has the top overall score among specialized models on OmniDocBench v1.6 and Wild_OmniDocBench, and the top Table TEDS on both. Its table lead doesn’t hold on every benchmark: MinerU 2.5 Pro scores higher on Dr.DocBench tables, and several models score higher on PureDocBench’s digital-degraded tables.
If formula accuracy matters most, compare before you choose. On OmniDocBench v1.6, OvisOCR2, PaddleOCR-VL-1.6, MinerU2.5-Pro, GLM-OCR and others score higher on Formula CDM. TeleOCR also trails several models on Wild_OmniDocBench formulas, but it leads on PureDocBench’s clean and digital-degraded formula scores. The card lists only Chinese and English, so if you need other languages, test them before you commit. TeleOCR also doesn’t fit if you can’t run trust_remote_code, or need a complete end-to-end page pipeline from the model card alone. For that pipeline, the card points to the GitHub repo.
Related
Parse Documents to Markdown with Jina-OCR-v1 on One GPU
Jina-OCR-v1 is a 570M-active-parameter OCR model that turns page images into Markdown. Here is how to run it with Transformers, vLLM, or the hosted API.
jinaai/jina-ocr-v1
Run Gemma 4 31B for Image, Video and Reasoning Tasks
Google DeepMind's largest dense Gemma 4 model handles images, video frames and 256K-token context. Here's how to run it with Transformers.
google/gemma-4-31B-it
DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash