Constarium
← Search

Data · dataset · 2026

A Tibetan OCR Dataset for Complex-Layout Documents(Ti-OCR)

Listed in ScienceDB

“A Tibetan OCR Dataset for Complex-Layout Documents” is constructed for Tibetan text detection, text recognition, and complex document layout understanding in real publication and paper-document scenarios.

Description

The data are mainly collected from books, newspapers, and various paper documents, covering body text, titles, tables of contents, notes, tables, formulas, text embedded in figures, publication information, and multilingual or multimodal mixed-layout elements.

The original images were captured using smartphones and digital cameras under natural light or indoor lighting conditions. No additional layout rectification or manual image enhancement was applied during data collection, so as to preserve common characteristics of real document images, including perspective distortion, slight page bending, uneven illumination, page shadows, show-through, edge cropping, table-line interference, and image-text mixed layouts.

Read the rest (5 more)

This dataset is a document image dataset and does not contain continuous temporal observations or geospatial observation variables; therefore, acquisition time and fixed geographic spatial coverage are not used as organizational dimensions of the dataset.The dataset has a total volume of 3.08 GB and consists of 1,080 JPG images and 1,080 corresponding JSON annotation files, with 41,098 annotated text instances in total.

According to page structure and layout complexity, the dataset is divided into four layout categories: full-text pages (FTS), table-text mixed pages (TTM), image-text mixed pages (TIM), and multi-element mixed pages (MEM). The numbers of samples in these four categories are 452, 99, 321, and 208, accounting for 41.85%, 9.17%, 29.72%, and 19.26%, respectively. At the instance level, the dataset provides five language or content-type labels, including Tibetan (TI), Chinese (CH), English (EN), digits (DI), and formulas (FM), with 34,818, 1,110, 1,292, 3,609, and 269 instances, respectively.

Each image contains 38.05 annotated text instances on average, with the number of instances per image ranging from 1 to 497. The dataset therefore covers multiple levels of page complexity, from simple text pages to dense tables, image-text mixed pages, and multi-element mixed layouts.This dataset can be used for Tibetan complex-layout document text detection, text recognition, end-to-end OCR, table text parsing, image-text mixed page analysis, multilingual document OCR, publication digitization, and low-resource language document intelligence research.

The data are organized in common JPG image and JSON annotation formats and can be read, parsed, converted, trained, and evaluated using Python together with tools such as OpenCV, Pillow, json, and PaddleOCR. For text detection tasks, the polygon coordinates in the points field can be directly used as supervision signals. For text recognition tasks, text regions can be cropped according to the annotated coordinates, and the text field can be used as the recognition label.

The samples are mainly derived from publicly available publications and various paper document images and do not involve personal sensitive information. The dataset is recommended for academic research, algorithm training, and performance evaluation, subject to the usage rules and citation requirements of the data publishing platform.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Image 75% · Linguistics 69% · Optical character recognition 65% · Tabular 65% · Text 75%
Provenance · 1 source records, 15 field assertions
SourceKeyLast seenRaw
ScienceDB10.57760/sciencedb.j00001.018769 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · scidb cnconnector:scidb_cn@1.0.0
concepts[field].anzsrc:group:4704enrichment · scidb cntaxonomy-embedding@1.0.0title+keywords+description (69%)
concepts[field].local:field:earth-environmentalmapping · scidb cnconnector:scidb_cn@1.0.0
concepts[field].local:field:engineeringmapping · scidb cnconnector:scidb_cn@1.0.0
concepts[field].local:field:humanitiesmapping · scidb cnconnector:scidb_cn@1.0.0
concepts[field].local:field:life-sciencesmapping · scidb cnconnector:scidb_cn@1.0.0
concepts[field].local:field:social-sciencemapping · scidb cnconnector:scidb_cn@1.0.0
concepts[method].local:method:ocrenrichment · scidb cnkeyword-concept-rules@1.0.0title+description (65%)
concepts[modality].local:modality:imageenrichment · scidb cnkeyword-concept-rules@1.0.0title+description (75%)
concepts[modality].local:modality:tabularenrichment · scidb cnkeyword-concept-rules@1.0.0title+description (65%)
concepts[modality].local:modality:textenrichment · scidb cnkeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · scidb cnconnector:scidb_cn@1.0.0/metadata/dc/description
licensesource · scidb cnconnector:scidb_cn@1.0.0/metadata/dc/rights
publication_datesource · scidb cnconnector:scidb_cn@1.0.0
titlesource · scidb cnconnector:scidb_cn@1.0.0/metadata/dc/title