Hugging Face Datasets2026 · Image
rustensai/russian-handwriting-ocr…Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста.… …1305 Уникальных текстов: 575 Средняя длина текста: 3790 символов Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/rustensai/russian-handwriting-ocr…
Hugging Face Datasets2024 · Image
Thai Handwriting Dataset…Thai Handwriting Dataset This dataset combines two major Thai handwriting datasets: BEST 2019 Thai Handwriting Recognition dataset (train-0000.parquet) Thai Handwritten Free Dataset…
Hugging Face Datasets2025 · Image
Boston Public Library Card CatalogBoston Public Library Rare Books Card Catalog Dataset Dataset Description This dataset contains approximately 410,000 digitized catalog cards from the Boston Public Library's Rare Books Department card catalog. The cards represent the main entry catalog (author/title cross-referenced) covering printed materials from various historical periods. Why This Dataset? Historical card catalogs are rich so
Hugging Face Datasets2026 · dataset
African Medical Records (AMR): Nigerian Handwritten Clinical Records…Why AMR Exists Clinical documentation across much of Africa is still handwritten, and the world's OCR and handwritten-text-recognition (HTR) systems have almost never seen it.… …Models trained on Western clinical forms or clean printed text fail on real Nigerian ward notes, prescriptions, and observation charts, where handwriting styles, abbreviations, drug…
Hugging Face Datasets2026 · Image
Burmese Handwritten Sentence Dataset (BHSD)…Burmese Handwritten Sentence Dataset (BHSD) BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR… …The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers. This dataset would not have been possible without its volunteers.…
Hugging Face Datasets2022 · Table · Parquet
Berlin State Library OCR…At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.… …For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012).…
Hugging Face Datasets2026 · dataset
OCR Synthetic Multilingual v1…OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition.… …This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset…
Hugging Face Datasets2026 · Image
sarvamai/indic-ocr-bench…Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor.… …The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.…
Hugging Face Datasets2026 · Table · Parquet
PubMed-OCR…PubMed-OCR: PMC Open Access OCR Annotations PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs.… …on the dataset page: https://huggingface.co/datasets/bevaya/pubmed-ocr.…
Hugging Face Datasets2026 · Table · CSV
Sintético de citações jurídicas (Desafio Jusbrasil BRACIS 2026)Sintético de citações jurídicas — Desafio Jusbrasil BRACIS 2026 Pareceres jurídicos sintéticos em português, com gabarito de cada citação (posição exata no texto e classe). Foram gerados pela equipe para medir e calibrar um verificador de citações. A regra do desafio exige que dado sintético usado em treino seja publicado, e este repositório cumpre isso. São duas versões, com o mesmo gabarito (mes
City of Seattle Open Data portal2024 · dataset
Discrimination Case Closures by Month and Type, 2017-PresentA closed discrimination case is an investigation that has been completed at the Seattle Office for Civil Rights (SOCR). This dataset shows closed cases by month and case type.
Hugging Face Datasets2025 · Image
MyanmarOCR-ImageText…🇲🇲 MyanmarOCR-ImageText Dataset A clean and diverse Burmese Image-to-Text dataset for OCR and multimodal AI research. 📌 Summary Total images: 41,664 Unique Burmese text entries:… …1,139 Styles per text: 32 variations each Resolution: 512 × 512 File types: PNG/JPG images Dataset split: train only Use cases: OCR, I2T (image-to-text), VLM pretrain/fine-tune All…
Hugging Face Datasets2025 · Image
IIIT5KMETA https://github.com/open-mmlab/mmocr/blob/main/dataset_zoo/iiit5k/metafile.yml Name: 'IIIT5K' Paper: Title: Scene Text Recognition using Higher Order Language Priors URL: http://cvit.iiit.ac.in/projects/SceneTextUnderstanding/Home/mishraBMVC12.pdf Venue: BMVC Year: '2012' BibTeX: '@InProceedings{MishraBMVC12, author = "Mishra, A. and Alahari, K. and Jawahar, C.~V.", title = "Scene Text Recogni
Hugging Face Datasets2022 · Image
Europeana NewspapersDataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiativ
Hugging Face Datasets2023 · dataset
HWR200…HWR200: New open access dataset of handwritten texts images in Russian This is a dataset of handwritten texts images in Russian created by 200 writers with different handwriting and…
Hugging Face Datasets2026 · Image
PDFA OCR Dataset - KREATIVE TIME BOX…PDFA OCR Dataset Curated and Published by KREATIVE TIME BOX This dataset contains document page images along with their corresponding OCR layout bounding box annotations derived from… …document page renders) Annotation Format: JSON with text lines, normalized bounding… See the full description on the dataset page: https://huggingface.co/datasets/infokreativetimebox/ktb-ocr-dataset…
Hugging Face Datasets2026 · Image
AdvSpot…AdvSpot AdvSpot is the first grounded adversarial OCR benchmark, introduced in the paper ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation…
Hugging Face Datasets2022 · Image
chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
Hugging Face Datasets2022 · dataset
ai-forever/Peter…The dataset annotation contain end-to-end markup for training detection and OCR models, as well as an end-to-end model for reading text from pages.…
Hugging Face Datasets2022 · dataset
ai-forever/school_notebooks_RU…The dataset annotation contain end-to-end markup for training detection and OCR models, as well as an end-to-end model for reading text from pages.…
City of Seattle Open Data portal2021 · dataset
City of Seattle Racial Equity ActionsThis dataset includes the actions the City of Seattle has been taking towards racial equity. We publish it for accountability and transparency. See who to contact, what we'll deliver, and how we plan on meeting our desired outcomes. You can also view this information here: https://www.seattle.gov/rsji/city-racial-equity-actions#/1
Hugging Face Datasets2022 · Image
ai-forever/school_notebooks_EN…The dataset annotation contain end-to-end markup for training detection and OCR models, as well as an end-to-end model for reading text from pages.…
Hugging Face Datasets2026 · Image
ParseBenchParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics d
Hugging Face Datasets2026 · Image
Ledger…Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values.…
INESC TEC Research Data Repository2022 · Table · CSV
Manual Transcriptions of Typewritten Digital Representations of Portuguese Cultural Heritage Documents from the 20th CenturyThe dataset includes manual transcriptions of typewritten digital representations of Portuguese cultural heritage documents from the 20th century, extracted from the Arquivo Nacional Torre do Tombo (ANTT) (https://digitarq.arquivos.pt/). It has records from two funds: the General Administration of National Treasury (DGFP) and the National Secretariat of Information (SNI). The manual transcriptions