Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 10 min read

Parse Scanned and Phone-Photo Documents with TeleOCR (1.2B)

TeleOCR is a 1.2B Apache-2.0 model that turns text, tables, formulas and layouts into structured output, including from photos of warped pages.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

TeleOCR is a ~1.2B-parameter vision-language model for document parsing, released under Apache-2.0. Its authors are Peng Cai and colleagues. The model was first published as NaviDC-OCR under StarDoc-AI. Most document parsers handle either clean digital pages or camera-captured pages. TeleOCR handles both with one model. According to the card, it parses layout and content on distorted documents without a separate dewarping or rectification step. The card’s evidence for this is visual examples on the DocUNet and DIR300 datasets, with no scores. One set of weights covers plain text, tables, LaTeX formulas, code, page layout and extracting the table implied by a scientific figure. You choose the task with a fixed prompt string.

Key specs

  • Size: ~1.2B parameters. The card doesn’t describe the architecture or backbone. Its acknowledgements list MinerU, Qwen2.5-VL, Qwen3, Transformers, PyTorch and FlashAttention.
  • Languages: the card’s metadata lists Chinese and English
  • Outputs: plain text, tables in OTSL format (the card includes an HTML converter), LaTeX formulas, code, layout analysis, and tables extracted from scientific figures
  • License: Apache-2.0

The authors report these benchmark numbers on the card. None of them come from independent evaluation:

Benchmark TeleOCR Closest competitor on the card
OmniDocBench v1.6 overall 96.87 OvisOCR2 (0.8B): 96.58
OmniDocBench v1.6 Table TEDS 97.05 OvisOCR2 / PaddleOCR-VL-1.6 / HunyuanOCR-1.5: 94.76
Wild_OmniDocBench overall 88.53 OvisOCR2: 87.91
PureDocBench real-degraded overall 70.85 Gemini-3.1-Pro: 71.98 (ahead of TeleOCR)
Dr.DocBench Challenge overall 67.96 MinerU 2.5 Pro: 62.26

In the card’s OmniDocBench v1.6 table, the HunyuanOCR-1.5 row matches the PaddleOCR-VL-1.6 row in every column except Overall (94.74 vs 96.33). The other five metrics are identical (0.033 / 97.49 / 94.76 / 97.11 / 0.127). That may be a copy error, so treat comparisons against either model with care.

Results differ by benchmark:

  • Tables. TeleOCR has the top Table TEDS on OmniDocBench v1.6 (97.05) and Wild_OmniDocBench (89.05). Other benchmarks show something different. On Dr.DocBench its Table TEDS is 64.97, behind MinerU 2.5 Pro at 67.75. On PureDocBench’s digital-degraded table score it gets 80.45, behind OvisOCR2 (84.71), FD-RL (83.22) and MinerU2.5-Pro (80.73).
  • Formulas. On OmniDocBench v1.6 its Formula CDM is 96.36. That is below OvisOCR2 (97.53), PaddleOCR-VL-1.6 and HunyuanOCR-1.5 (97.49), MinerU2.5-Pro (97.45), GLM-OCR (97.18) and PaddleOCR-VL-1.5 (96.69). On Wild_OmniDocBench, five listed models score higher than its 88.26. PureDocBench is different: TeleOCR has the top formula score on the clean and digital-degraded splits. On the real-degraded split, Gemini-3.1-Pro (68.62) beats it (65.11).
  • Real-degraded photos. On PureDocBench’s real-degraded split, Gemini-3.1-Pro has the higher overall score.

Install

Terminal window
pip install transformers torch pillow

Run it

The card loads the model through AutoModel with trust_remote_code=True and wraps generation in an infer() helper. The examples below import this file. It loads the model at import time, so each script that imports it loads the model once when it starts.

teleocr.py
import torch
from PIL import Image
from transformers import AutoProcessor, AutoModel
# The id used in the card's own code. The card itself is hosted at XingChen-AGI/TeleOCR.
MODEL_ID = "StarDoc-AI/TeleOCR"
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True, use_fast=True)
model = AutoModel.from_pretrained(
MODEL_ID,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).cuda().eval()
PROMPT_TEXT = "Please output the text content from the image."
PROMPT_TABLE = "This is the image of a table. Please output the table in OTSL format."
PROMPT_FORMULA = "Please write out the expression of the formula in the image using LaTeX format."
PROMPT_CODE = "The image contains a code snippet, please output the parsing result."
PROMPT_LAYOUT = "Analyze the image layout."
PROMPT_LAYOUT_DISTORTED = "\nMulti-point Layout Segmentation Analysis."
PROMPT_FIGURE = "This is a scientific figure. Please extract the table implied by this figure."
def infer(image: Image.Image, prompt: str) -> str:
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [
{"type": "image"},
{"type": "text", "text": prompt},
]},
]
chat_prompt = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = processor(
text=[chat_prompt],
images=[image.convert("RGB")],
padding=True,
return_tensors="pt",
).to(device=model.device, dtype=model.dtype)
output_ids = model.generate(
**inputs,
use_cache=True,
max_new_tokens=4096,
do_sample=False,
)
output_ids = output_ids.cpu().tolist()[0][len(inputs.input_ids[0]):]
return processor.batch_decode(
[output_ids],
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0].strip()
if __name__ == "__main__":
import sys
image = Image.open(sys.argv[1]).convert("RGB")
print(infer(image, PROMPT_TEXT))
Terminal window
python teleocr.py page.png

Run the text prompt over a folder of images

This script runs the text prompt on every image in a folder and saves one .txt file per image. It processes one image at a time, as the card does. The card only demonstrates the text prompt on a single sample image. It doesn’t show results on full pages, and it points to the GitHub repo for complete document parsing. Check the output on your own pages before you rely on it.

batch_text.py
import sys
from pathlib import Path
from PIL import Image
from teleocr import infer, PROMPT_TEXT
src = Path(sys.argv[1])
dst = Path(sys.argv[2])
dst.mkdir(parents=True, exist_ok=True)
for path in sorted(src.iterdir()):
if path.suffix.lower() not in {".png", ".jpg", ".jpeg"}:
continue
image = Image.open(path).convert("RGB")
text = infer(image, PROMPT_TEXT)
(dst / f"{path.stem}.txt").write_text(text, encoding="utf-8")
print(f"{path.name}: {len(text)} chars")
Terminal window
python batch_text.py ./scans ./out

Extract tables to HTML

TeleOCR scores highest on tables in OmniDocBench v1.6 (97.05 TEDS) and Wild_OmniDocBench (89.05 TEDS). As shown above, other benchmarks on the card rank it lower. The model outputs tables in OTSL, a compact token format, not HTML. The OTSL-to-HTML converter below is copied from the card. The card also runs it on the output of the scientific-figure prompt, which asks for the table implied by a figure.

otsl.py
import html
import itertools
import re
from dataclasses import dataclass
@dataclass
class TableCell:
text: str
start_row_offset_idx: int
end_row_offset_idx: int
start_col_offset_idx: int
end_col_offset_idx: int
row_span: int = 1
col_span: int = 1
OTSL_NL = "<nl>"
OTSL_FCEL = "<fcel>"
OTSL_ECEL = "<ecel>"
OTSL_LCEL = "<lcel>"
OTSL_UCEL = "<ucel>"
OTSL_XCEL = "<xcel>"
OTSL_TOKENS = [OTSL_NL, OTSL_FCEL, OTSL_ECEL, OTSL_LCEL, OTSL_UCEL, OTSL_XCEL]
def _otsl_extract_tokens_and_text(text: str):
pattern = "(" + "|".join(map(re.escape, OTSL_TOKENS)) + ")"
tokens = re.findall(pattern, text)
parts = [part for part in re.split(pattern, text) if part.strip()]
return tokens, parts
def _count_right(rows, row_idx, col_idx, tokens):
span = 0
while col_idx < len(rows[row_idx]) and rows[row_idx][col_idx] in tokens:
span += 1
col_idx += 1
return span
def _count_down(rows, row_idx, col_idx, tokens):
span = 0
while row_idx < len(rows) and col_idx < len(rows[row_idx]) and rows[row_idx][col_idx] in tokens:
span += 1
row_idx += 1
return span
def _otsl_parse_texts(parts, tokens):
rows = [list(row) for is_nl, row in itertools.groupby(tokens, lambda token: token == OTSL_NL) if not is_nl]
if not rows:
return [], []
max_cols = max(len(row) for row in rows)
for row in rows:
row.extend([OTSL_ECEL] * (max_cols - len(row)))
cells = []
row_idx = 0
col_idx = 0
for idx, part in enumerate(parts):
if part in (OTSL_FCEL, OTSL_ECEL):
cell_text = ""
right_offset = 1
if part != OTSL_ECEL and idx + 1 < len(parts) and parts[idx + 1] not in OTSL_TOKENS:
cell_text = parts[idx + 1].strip()
right_offset = 2
next_right = parts[idx + right_offset] if idx + right_offset < len(parts) else ""
next_bottom = rows[row_idx + 1][col_idx] if row_idx + 1 < len(rows) and col_idx < len(rows[row_idx + 1]) else ""
col_span = 1 + (_count_right(rows, row_idx, col_idx + 1, {OTSL_LCEL, OTSL_XCEL}) if next_right in {OTSL_LCEL, OTSL_XCEL} else 0)
row_span = 1 + (_count_down(rows, row_idx + 1, col_idx, {OTSL_UCEL, OTSL_XCEL}) if next_bottom in {OTSL_UCEL, OTSL_XCEL} else 0)
cells.append(TableCell(
text=cell_text,
row_span=row_span,
col_span=col_span,
start_row_offset_idx=row_idx,
end_row_offset_idx=row_idx + row_span,
start_col_offset_idx=col_idx,
end_col_offset_idx=col_idx + col_span,
))
if part in (OTSL_FCEL, OTSL_ECEL, OTSL_LCEL, OTSL_UCEL, OTSL_XCEL):
col_idx += 1
elif part == OTSL_NL:
row_idx += 1
col_idx = 0
return cells, rows
def convert_otsl_to_html(otsl_content: str) -> str:
if otsl_content.startswith("<table") and otsl_content.endswith("</table>"):
return otsl_content
tokens, parts = _otsl_extract_tokens_and_text(otsl_content)
cells, rows = _otsl_parse_texts(parts, tokens)
if not cells or not rows:
return ""
grid = [[None for _ in range(len(rows[0]))] for _ in range(len(rows))]
for cell in cells:
for row_idx in range(cell.start_row_offset_idx, min(cell.end_row_offset_idx, len(rows))):
for col_idx in range(cell.start_col_offset_idx, min(cell.end_col_offset_idx, len(rows[0]))):
grid[row_idx][col_idx] = cell
html_rows = []
for row_idx, row in enumerate(grid):
html_rows.append("<tr>")
for col_idx, cell in enumerate(row):
if cell is None or cell.start_row_offset_idx != row_idx or cell.start_col_offset_idx != col_idx:
continue
attrs = ""
if cell.row_span > 1:
attrs += f' rowspan="{cell.row_span}"'
if cell.col_span > 1:
attrs += f' colspan="{cell.col_span}"'
html_rows.append(f"<td{attrs}>{html.escape(cell.text.strip())}</td>")
html_rows.append("</tr>")
return "<table>" + "".join(html_rows) + "</table>"
tables.py
import sys
from PIL import Image
from otsl import convert_otsl_to_html
from teleocr import infer, PROMPT_TABLE, PROMPT_FIGURE
PROMPTS = {"table": PROMPT_TABLE, "figure": PROMPT_FIGURE}
mode, path = sys.argv[1], sys.argv[2]
if mode not in PROMPTS:
sys.exit(f"Unknown mode {mode!r}: use 'table' or 'figure'")
prompt = PROMPTS[mode]
image = Image.open(path).convert("RGB")
raw_otsl = infer(image, prompt)
table_html = convert_otsl_to_html(raw_otsl)
if not table_html:
print("Could not parse OTSL output:\n" + raw_otsl)
else:
print(table_html)
Terminal window
python tables.py table invoice_table.png > table.html
python tables.py figure scientific_figure.png > figure_data.html

Formulas to LaTeX

The formula prompt returns LaTeX. The card passes that output through a post_process step that strips \[ ... \] and wraps the result in $$...$$. The ContentBlock dataclass and post_process function below are copied from the card. The formula call follows the card’s example, which wraps the whole image in one equation block. Keep in mind that on the card’s OmniDocBench v1.6 and Wild_OmniDocBench results, several models score higher than TeleOCR on formulas.

formula.py
import sys
from dataclasses import dataclass
from PIL import Image
from otsl import convert_otsl_to_html
from teleocr import infer, PROMPT_FORMULA
@dataclass
class ContentBlock:
type: str
bbox: list[float]
angle: int | None = None
content: str | None = None
def post_process(blocks: list[ContentBlock]) -> list[ContentBlock]:
for block in blocks:
if block.type == "table" and block.content:
block.content = convert_otsl_to_html(block.content)
elif block.type == "equation" and block.content:
content = block.content.strip()
content = content.removeprefix("\\[").removesuffix("\\]").strip()
if not (content.startswith("$") and content.endswith("$")):
content = f"$${content}$$"
block.content = content
return [block for block in blocks if block.type != "equation_block"]
if __name__ == "__main__":
image = Image.open(sys.argv[1]).convert("RGB")
raw_formula = infer(image, PROMPT_FORMULA)
formula_block = ContentBlock("equation", [0.0, 0.0, 1.0, 1.0], content=raw_formula)
print(post_process([formula_block])[0].content)
Terminal window
python formula.py formula.png

Layout analysis on distorted pages

The card has two layout prompts: one for regular pages and one for distorted documents. The distorted-document prompt starts with a newline. Before running either one, the card resizes the image to 1036×1036 with bicubic resampling. The card supports its distorted-document claims with visual examples on DocUNet and DIR300, not scores. The card also doesn’t document the layout output format, so this example prints the raw result. Check it on your own data before you write a parser for it.

layout.py
import sys
from PIL import Image
from teleocr import infer, PROMPT_LAYOUT, PROMPT_LAYOUT_DISTORTED
path = sys.argv[1]
distorted = len(sys.argv) > 2 and sys.argv[2] == "--distorted"
image = Image.open(path).convert("RGB")
image = image.resize((1036, 1036), Image.Resampling.BICUBIC)
prompt = PROMPT_LAYOUT_DISTORTED if distorted else PROMPT_LAYOUT
print(infer(image, prompt).strip())
Terminal window
python layout.py scanned_page.jpg
python layout.py phone_photo.jpg --distorted

Gotchas

  • Repo id mismatch. The card is hosted at XingChen-AGI/TeleOCR, but its code loads StarDoc-AI/TeleOCR, and that is the id these examples use. The model was also renamed from NaviDC-OCR. We haven’t checked which repo id loads with trust_remote_code. If one doesn’t work, try the other.
  • trust_remote_code=True is required. The model runs custom code from the repo, so review that code before you deploy it.
  • The example assumes a CUDA GPU in bf16. It calls .cuda() and casts inputs to model.dtype. The card doesn’t document any other hardware path. A community GGUF conversion with llama.cpp support exists. It was made from NaviDC-OCR before the rename, and the card doesn’t say whether it matches the current TeleOCR weights or accepts the same prompts.
  • Use the exact prompts. Tasks are selected by the fixed prompt strings shown above. The distorted-layout prompt starts with \n, so keep it.
  • Resize images for layout. The card resizes to 1036×1036 for both layout prompts.
  • Tables come out as OTSL, not HTML. Run the output through convert_otsl_to_html. Formulas need the card’s post_process step (included in formula.py above) if you want $$...$$ delimiters.
  • max_new_tokens=4096. Very dense inputs may get cut off at this limit.
  • Model loads on import. teleocr.py loads the weights when it’s imported. Each script above loads the model once per run.
  • Single-region tasks. The card shows one task per image. For complete document parsing, the card points to the GitHub repo and doesn’t describe the pipeline itself.
  • Languages. The card’s metadata lists only Chinese (zh) and English (en). It doesn’t say anything about other languages, so test them yourself if you need them.
  • Self-reported benchmarks. Every score in this post comes from the authors’ card.

When to pick it

Pick TeleOCR if you need a small, Apache-2.0 parser for Chinese or English documents, especially if some input arrives as camera photos instead of clean PDFs. On the card it has the top overall score among specialized models on OmniDocBench v1.6 and Wild_OmniDocBench, and the top Table TEDS on both. Its table lead doesn’t hold on every benchmark: MinerU 2.5 Pro scores higher on Dr.DocBench tables, and several models score higher on PureDocBench’s digital-degraded tables.

If formula accuracy matters most, compare before you choose. On OmniDocBench v1.6, OvisOCR2, PaddleOCR-VL-1.6, MinerU2.5-Pro, GLM-OCR and others score higher on Formula CDM. TeleOCR also trails several models on Wild_OmniDocBench formulas, but it leads on PureDocBench’s clean and digital-degraded formula scores. The card lists only Chinese and English, so if you need other languages, test them before you commit. TeleOCR also doesn’t fit if you can’t run trust_remote_code, or need a complete end-to-end page pipeline from the model card alone. For that pipeline, the card points to the GitHub repo.

Related