Hugging Face Datasets2026 · Image
rustensai/russian-handwriting-ocr…Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста.… …1305 Уникальных текстов: 575 Средняя длина текста: 3790 символов Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/rustensai/russian-handwriting-ocr…
Hugging Face Datasets2026 · Image · gated
5CD-AI/Viet-Handwriting-OCR-v2WE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2024 · Image
Thai Handwriting Dataset…Thai Handwriting Dataset This dataset combines two major Thai handwriting datasets: BEST 2019 Thai Handwriting Recognition dataset (train-0000.parquet) Thai Handwritten Free Dataset…
Hugging Face Datasets2026 · Image · gated
Handwritten GCSE Exam Answers Dataset…Usage Examples Basic Dataset Loading from datasets import load_dataset # Load the full dataset dataset = load_dataset("JunaidMB/handwriting-ocr-images-dataset") # Access specific splits… …train_data = dataset['train'] # 62 samples test_data =… See the full description on the dataset page: https://huggingface.co/datasets/JunaidMB/handwriting-ocr-images-dataset.…
Hugging Face Datasets2026 · Image · gated
5CD-AI/VietHTR-LineWE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2025 · Image
Boston Public Library Card CatalogBoston Public Library Rare Books Card Catalog Dataset Dataset Description This dataset contains approximately 410,000 digitized catalog cards from the Boston Public Library's Rare Books Department card catalog. The cards represent the main entry catalog (author/title cross-referenced) covering printed materials from various historical periods. Why This Dataset? Historical card catalogs are rich so
Hugging Face Datasets2026 · Image
Burmese Handwritten Sentence Dataset (BHSD)…Burmese Handwritten Sentence Dataset (BHSD) BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR… …The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers. This dataset would not have been possible without its volunteers.…
Hugging Face Datasets2026 · Image
sarvamai/indic-ocr-bench…Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor.… …The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.…
Hugging Face Datasets2025 · Image
MyanmarOCR-ImageText…🇲🇲 MyanmarOCR-ImageText Dataset A clean and diverse Burmese Image-to-Text dataset for OCR and multimodal AI research. 📌 Summary Total images: 41,664 Unique Burmese text entries:… …1,139 Styles per text: 32 variations each Resolution: 512 × 512 File types: PNG/JPG images Dataset split: train only Use cases: OCR, I2T (image-to-text), VLM pretrain/fine-tune All…
Hugging Face Datasets2025 · Image
IIIT5KMETA https://github.com/open-mmlab/mmocr/blob/main/dataset_zoo/iiit5k/metafile.yml Name: 'IIIT5K' Paper: Title: Scene Text Recognition using Higher Order Language Priors URL: http://cvit.iiit.ac.in/projects/SceneTextUnderstanding/Home/mishraBMVC12.pdf Venue: BMVC Year: '2012' BibTeX: '@InProceedings{MishraBMVC12, author = "Mishra, A. and Alahari, K. and Jawahar, C.~V.", title = "Scene Text Recogni
Hugging Face Datasets2022 · Image
Europeana NewspapersDataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiativ
Hugging Face Datasets2026 · Image
PDFA OCR Dataset - KREATIVE TIME BOX…PDFA OCR Dataset Curated and Published by KREATIVE TIME BOX This dataset contains document page images along with their corresponding OCR layout bounding box annotations derived from… …document page renders) Annotation Format: JSON with text lines, normalized bounding… See the full description on the dataset page: https://huggingface.co/datasets/infokreativetimebox/ktb-ocr-dataset…
Hugging Face Datasets2026 · Image
AdvSpot…AdvSpot AdvSpot is the first grounded adversarial OCR benchmark, introduced in the paper ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation…
Hugging Face Datasets2022 · Image
chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
Hugging Face Datasets2022 · Image
ai-forever/school_notebooks_EN…The dataset annotation contain end-to-end markup for training detection and OCR models, as well as an end-to-end model for reading text from pages.…
Hugging Face Datasets2026 · Image
ParseBenchParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics d
Hugging Face Datasets2026 · Image · gated
Institutional Newspapers: Boston Public Library…Public Library. 1,473,635 public domain newspaper scans, published between 1795 and 1930 83,147,041 individual crops segmented from those scans 16.3 billion o200k_base tokens of VLM OCR…
Hugging Face Datasets2026 · Image
Ledger…Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values.…
Hugging Face Datasets2026 · Image
krotreaksmey/khmer-math-textbook…🇰🇭 Khmer Math Textbook Line-Level OCR Dataset A large-scale, high-resolution dataset of line-level Khmer text and mathematical formulas extracted from official Cambodian Grade 9,…
Hugging Face Datasets2026 · Image
Synthetic Medical Document Recognition BenchmarkSynthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the d
Hugging Face Datasets2026 · Image
Indonesian KTP Dataset 24K (Flat & 3D Perspective Augmented)…synthetic dataset designed to advance SOTA (State-of-the-Art) research in Document Information Extraction (DIE), Key Information Extraction (KIE), and Optical Character Recognition (OCR…
Hugging Face Datasets2026 · Image
Synthetic Text Images (English)Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise.
Hugging Face Datasets2026 · Image
DuckerMaster/Thai-Synth-Receipts…large-scale, highly robust synthetic dataset of Thai commercial documents designed specifically for training and evaluating state-of-the-art Document AI and Optical Character Recognition (OCR…
Hugging Face Datasets2026 · Image
PureDocBenchMain Leaderboard 58 models · 3 matched tracks · 🏆 Search, filter & sort the leaderboard → The current evaluation covers 13 pipeline / multi-stage specialists, 19 end-to-end specialists, and 26 general-purpose VLMs. Each track contains 1,475 pages. The top 10 by the three-track mean, Avg₃, are shown below. Rank Model (release) Type Clean ↑ Digital ↑ Real ↑ Avg₃ ↑ 1 GLM-5.3-Flash (2026-08) General V
Hugging Face Datasets2026 · Image · gated
OCRGenBench…OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities 🔐 Dataset Access This dataset is gated.… …You will receive an email confirmation once access is granted. 📖 Overview OCRGenBench is the most comprehensive benchmark to date for evaluating the OCR generative… See the full description…