Hugging Face Datasets2026 · Image
rustensai/russian-handwriting-ocr…Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста.… …1305 Уникальных текстов: 575 Средняя длина текста: 3790 символов Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/rustensai/russian-handwriting-ocr…
Hugging Face Datasets2026 · Image · gated
5CD-AI/Viet-Handwriting-OCR-v2WE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2026 · Image · gated
Handwritten GCSE Exam Answers Dataset…Usage Examples Basic Dataset Loading from datasets import load_dataset # Load the full dataset dataset = load_dataset("JunaidMB/handwriting-ocr-images-dataset") # Access specific splits… …train_data = dataset['train'] # 62 samples test_data =… See the full description on the dataset page: https://huggingface.co/datasets/JunaidMB/handwriting-ocr-images-dataset.…
Hugging Face Datasets2026 · Image · gated
5CD-AI/VietHTR-LineWE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2026 · dataset
African Medical Records (AMR): Nigerian Handwritten Clinical Records…Why AMR Exists Clinical documentation across much of Africa is still handwritten, and the world's OCR and handwritten-text-recognition (HTR) systems have almost never seen it.… …Models trained on Western clinical forms or clean printed text fail on real Nigerian ward notes, prescriptions, and observation charts, where handwriting styles, abbreviations, drug…
Hugging Face Datasets2026 · Image
Burmese Handwritten Sentence Dataset (BHSD)…Burmese Handwritten Sentence Dataset (BHSD) BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR… …The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers. This dataset would not have been possible without its volunteers.…
Hugging Face Datasets2026 · dataset
OCR Synthetic Multilingual v1…OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition.… …This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset…
e-cienciaDatos2026 · dataset · unknown
Minitutorial de OCR con un VLM formato taller: fine-tuning, inferencia y métricas con Gemma 3 4B (ipynb + instrucciones)…El presente recurso se trata de un minitutorial en formato taller que pretende ser una introducción a realizar OCR con un VLM (Vision Language Model) empleando Python, el repositorio… …No se necesitan conocimientos previos, ya que las celdas vienen preparadas para funcionar tras ejecutarse, pero se valoran conocimientos básicos de OCR y Python para sacar el máximo…
IISH Dataverse2026 · dataset · unknown
GLOBALISE Ground Truth for Handwritten Text and Layout RecognitionThis dataset contains Ground Truth PageXML files that were used to finetune the GLOBALISE Handwritten Text Recognition, baseline detection and region detection models (see Related Publications).
Hugging Face Datasets2026 · Image
sarvamai/indic-ocr-bench…Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor.… …The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.…
Hugging Face Datasets2026 · Table · Parquet
PubMed-OCR…PubMed-OCR: PMC Open Access OCR Annotations PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs.… …on the dataset page: https://huggingface.co/datasets/bevaya/pubmed-ocr.…
Hugging Face Datasets2026 · Table · CSV
Sintético de citações jurídicas (Desafio Jusbrasil BRACIS 2026)Sintético de citações jurídicas — Desafio Jusbrasil BRACIS 2026 Pareceres jurídicos sintéticos em português, com gabarito de cada citação (posição exata no texto e classe). Foram gerados pela equipe para medir e calibrar um verificador de citações. A regra do desafio exige que dado sintético usado em treino seja publicado, e este repositório cumpre isso. São duas versões, com o mesmo gabarito (mes
Hugging Face Datasets2026 · Image
PDFA OCR Dataset - KREATIVE TIME BOX…PDFA OCR Dataset Curated and Published by KREATIVE TIME BOX This dataset contains document page images along with their corresponding OCR layout bounding box annotations derived from… …document page renders) Annotation Format: JSON with text lines, normalized bounding… See the full description on the dataset page: https://huggingface.co/datasets/infokreativetimebox/ktb-ocr-dataset…
Hugging Face Datasets2026 · Image
AdvSpot…AdvSpot AdvSpot is the first grounded adversarial OCR benchmark, introduced in the paper ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation…
Hugging Face Datasets2026 · Image
ParseBenchParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics d
Hugging Face Datasets2026 · Image · gated
Institutional Newspapers: Boston Public Library…Public Library. 1,473,635 public domain newspaper scans, published between 1795 and 1930 83,147,041 individual crops segmented from those scans 16.3 billion o200k_base tokens of VLM OCR…
Hugging Face Datasets2026 · Image
Ledger…Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values.…
Hugging Face Datasets2026 · Image
krotreaksmey/khmer-math-textbook…🇰🇭 Khmer Math Textbook Line-Level OCR Dataset A large-scale, high-resolution dataset of line-level Khmer text and mathematical formulas extracted from official Cambodian Grade 9,…
IISH Dataverse2026 · dataset · unknown
GLOBALISE Laypa Region Model - August 2023This is a Laypa region detection model that was created to detect and identify regions (such as page number, heading, paragraph, and marginalia) on the scans of the GLOBALISE VOC corpus. It was trained on Ground Truth that is also available in this Dataverse (GLOBALISE Ground Truth for Handwritten Text and Layout Recognition) and applied using the Loghi Handwritten Text Recognition tools to genera
IISH Dataverse2026 · dataset · unknown
GLOBALISE Loghi Handwritten Text Recognition Model – August 2023This is a Loghi Handwritten Text Recognition (HTR) model that was created to transcribe text on the scans of the GLOBALISE VOC corpus. It was finetuned using Ground Truth that is also available in this Dataverse (GLOBALISE Ground Truth for Handwritten Text and Layout Recognition) and applied using the Loghi HTR tools to generate VOC transcriptions v2 - GLOBALISE. The model's hyperparameters are av
Hugging Face Datasets2026 · Text
Manga109-s Text Line AnnotationsManga109-s Text Line Annotations High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP). Notice: This d
Hugging Face Datasets2026 · Image
Synthetic Medical Document Recognition BenchmarkSynthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the d
Hugging Face Datasets2026 · Text
VotingBookletsVotingBooklets Dataset Summary VotingBooklets is a large-scale four-language parallel corpus extracted from the complete collection of Swiss federal voting booklets (Abstimmungsbüchlein), covering federal votes from June 1977 to March 2026. It contains aligned paragraph-level segments across German (de), French (fr), Italian (it), and Romansh Grischun (rm), and serves as a resource for low-resourc
Hugging Face Datasets2026 · Image
Indonesian KTP Dataset 24K (Flat & 3D Perspective Augmented)…synthetic dataset designed to advance SOTA (State-of-the-Art) research in Document Information Extraction (DIE), Key Information Extraction (KIE), and Optical Character Recognition (OCR…
Hugging Face Datasets2026 · Image
Synthetic Text Images (English)Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise.