Hugging Face Datasets2026 · dataset
African Medical Records (AMR): Nigerian Handwritten Clinical Records…Why AMR Exists Clinical documentation across much of Africa is still handwritten, and the world's OCR and handwritten-text-recognition (HTR) systems have almost never seen it.… …Models trained on Western clinical forms or clean printed text fail on real Nigerian ward notes, prescriptions, and observation charts, where handwriting styles, abbreviations, drug…
Hugging Face Datasets2026 · Image
Burmese Handwritten Sentence Dataset (BHSD)…Burmese Handwritten Sentence Dataset (BHSD) BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR… …The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers. This dataset would not have been possible without its volunteers.…
Hugging Face Datasets2022 · Table · Parquet
Berlin State Library OCR…At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.… …For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012).…
Hugging Face Datasets2026 · dataset
OCR Synthetic Multilingual v1…OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition.… …This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset…
Hugging Face Datasets2025 · Image
MyanmarOCR-ImageText…🇲🇲 MyanmarOCR-ImageText Dataset A clean and diverse Burmese Image-to-Text dataset for OCR and multimodal AI research. 📌 Summary Total images: 41,664 Unique Burmese text entries:… …1,139 Styles per text: 32 variations each Resolution: 512 × 512 File types: PNG/JPG images Dataset split: train only Use cases: OCR, I2T (image-to-text), VLM pretrain/fine-tune All…
Hugging Face Datasets2026 · Image
Ledger…Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values.…
Hugging Face Datasets2026 · Image
Synthetic Medical Document Recognition BenchmarkSynthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the d
Teesside University Research Data Repository2023 · dataset
Handwritten Numbers (HN)Handwritten numbers from one to three digits extracted from public documents. There are 10000 images of different sizes of numbers between 0 and 303 with a imbalanced distribution, there are even numbers without samples. This dataset was created to do proofs with classification models reuse. The images are named with numbers from 0 to 9999, the file labels.csv contain the labels. The number in the
Hugging Face Datasets2026 · Text
VotingBookletsVotingBooklets Dataset Summary VotingBooklets is a large-scale four-language parallel corpus extracted from the complete collection of Swiss federal voting booklets (Abstimmungsbüchlein), covering federal votes from June 1977 to March 2026. It contains aligned paragraph-level segments across German (de), French (fr), Italian (it), and Romansh Grischun (rm), and serves as a resource for low-resourc
Hugging Face Datasets2026 · Image
Indonesian KTP Dataset 24K (Flat & 3D Perspective Augmented)…synthetic dataset designed to advance SOTA (State-of-the-Art) research in Document Information Extraction (DIE), Key Information Extraction (KIE), and Optical Character Recognition (OCR…
Hugging Face Datasets2026 · Image
PureDocBenchMain Leaderboard 58 models · 3 matched tracks · 🏆 Search, filter & sort the leaderboard → The current evaluation covers 13 pipeline / multi-stage specialists, 19 end-to-end specialists, and 26 general-purpose VLMs. Each track contains 1,475 pages. The top 10 by the three-track mean, Avg₃, are shown below. Rank Model (release) Type Clean ↑ Digital ↑ Real ↑ Avg₃ ↑ 1 GLM-5.3-Flash (2026-08) General V
Hugging Face Datasets2026 · Table · Parquet
artefactory/ledger-long-context-KPI-QALEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TRE
Teesside University Research Data Repository2024 · dataset
BengaliPrintDB: A Repository of Machine-Printed Bengali Documents…Distinguishing between handwritten and machine-printed documents is vital in OCR applications due to varying processing methods.… …Training data selection differs, with diverse datasets for handwritten OCR models and uniform fonts for machine-printed OCR models.…
figshare + Loughborough Research Repository2026 · Astronomical catalogue
LongHisDoc: A Comprehensive Benchmark for Chinese Long Historical Document Understanding…/LongHisDoc_IMG), their OCR results (./OCR_res), and QA pairs (./LongHisDoc.json).</p>…
Teesside University Research Data Repository2026 · dataset
Multilingual AI-Based Dyslexia Detection…This dataset contains 852 character-level handwriting images collected from 100 unique children aged 6–10 years in school settings in Narowal, Punjab, Pakistan, for research on handwriting-based… …Urdu handwriting images were retained without OCR-based correction or linguistic normalization to preserve visually relevant characteristics such as dots, curves, loops, stroke formation…