AGP Picks
View all

China Telecom Unveils TeleOCR: Lightweight 1.2B Model Tops Global Document Parsing Benchmarks

Picture1

Open-sourced by China Telecom's Xingchen AGI Lab, the model tops three international benchmarks and wins the ICDAR 2026 competition, showing that a lightweight model can surpass general-purpose AI giants at specialized document understanding.

BEIJING, Sept. 30, 2026 (GLOBE NEWSWIRE) -- In recent days, China Telecom Artificial Intelligence Technology Co., Ltd. (China Telecom AI) announced that TeleOCR, its self-developed document parsing model, has set a new state-of-the-art (SOTA) result on OmniDocBench v1.6, the field's most comprehensive benchmark, with an overall score of 96.87 out of 100. The model also ranked first on two other benchmarks and took first place in the ICDAR 2026 Sci-ImageMiner Challenge, an international competition on scientific figure understanding.

With roughly 1.2 billion parameters, TeleOCR outperformed significantly larger specialized models — including MinerU 2.5-Pro and PaddleOCR-VL-1.6 — as well as general-purpose models such as Gemini 3 Pro and GPT-5.2 on document parsing tasks. Its code and weights are open-sourced on GitHub and Hugging Face.

"TeleOCR proves that precision engineering and targeted training can trump sheer scale," said a spokesperson for the Xingchen AGI Lab. "We built a 1.2-billion-parameter model that beats models dozens of times its size in document parsing, and we are sharing it with the global community."

Key achievements:

  • OmniDocBench v1.6: An overall score of 96.87 — the highest of all evaluated models — across 10 document types, 11 layouts, and five languages. TeleOCR ranked first in text recognition, table reconstruction, and reading-order restoration.
  • Wild-OmniDocBench v1.5: An overall score of 88.53 on camera-captured documents, about one point ahead of the runner-up and 4 to 10 points ahead of most end-to-end models.
  • PureDocBench: An average of 78.41 across three tracks, including a four-point lead over the second-place model on the most challenging "real degradation" track.
  • ICDAR 2026 Sci-ImageMiner Challenge: First place in the scientific-figure-to-table task, with a TEDS score more than two percentage points above the runner-up's.

TeleOCR addresses a persistent challenge in document AI: existing systems typically excel at either digital documents or camera-captured documents, but not both. Pipeline-based approaches handle clean digital files well but struggle with geometric distortion found in photographs. End-to-end models are more robust against distortion but fall short on high-resolution structured content such as tables and formulas.

TeleOCR resolves this trade-off through three innovations integrated into a unified framework. First, deformation-aware learning embeds geometric perception directly into the vision-language model, eliminating the need for external dewarping modules. Second, a content-structure decoupling strategy separates structural reasoning from content generation, enabling precise reconstruction of tables, formulas, and scientific charts. Third, a multi-model consensus voting mechanism generates high-quality training labels by aggregating predictions from heterogeneous models, avoiding the systematic bias inherent in single-model labeling.

"Most document parsing systems treat geometric correction as a separate preprocessing step. We turned it into an intrinsic capability of the model itself," the spokesperson said. "This means a photo of a wrinkled contract or a tilted whiteboard can be parsed accurately in one pass — no external plugins, no dewarping pipeline."

The model supports eight document parsing tasks, including digital layout detection, camera-captured layout segmentation, text recognition, formula recognition (LaTeX output), table recognition (OTSL/HTML output), code block recognition, scientific figure analysis, and seal recognition.

TeleOCR is available for immediate use through three channels: open-source code and model weights on GitHub and Hugging Face, and a production-ready API on China Telecom's Tianyi AI Open Platform.

The research team said TeleOCR's next phase will focus on deeper integration into high-value enterprise scenarios, including financial document processing, medical record digitization, academic research workflows and government archives — areas where converting unstructured documents into structured, machine-readable data can drive significant operational efficiency.

"This is about making documents truly machine-readable," the spokesperson said. "When a model can parse a photographed contract as accurately as a clean PDF, document automation becomes practical at scale."

About TeleOCR
TeleOCR is an open-source document parsing model developed by China Telecom's Xingchen AGI Lab. With approximately 1.2 billion parameters, it unifies the parsing of digital and camera-captured documents into structured outputs including Markdown, tables, and formulas. The model's code, weights, and technical paper are publicly available.

Download:
GitHub: https://github.com/caipeng328/TeleOCR
Hugging Face: https://huggingface.co/StarDoc-AI/TeleOCR

Contact
Xing Chen
DataAiTech-service@chinatelecom.cn

A photo accompanying this announcement is available at https://www.globenewswire.com/NewsRoom/AttachmentNg/a9e435f3-e459-4855-9e3f-1ed770a02d42


China Telecom AI

China Telecom AI

Legal Disclaimer:

EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

New Products Launch Guide

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.