Hugging Face Datasets2026 · Image
DuckerMaster/Thai-Synth-ReceiptsThai-Synth-Receipts Thai-Synth-Receipts is a large-scale, highly robust synthetic dataset of Thai commercial documents designed specifically for training and evaluating state-of-the-art Document AI and Optical Character Recognition (OCR) models. The dataset consists of 14,976 high-resolution document images (Receipts, Tax Invoices, Thermal Slips, and Quotations) across three distinct degradation v
Hugging Face Datasets2026 · Image
PESTPEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation 🎉 ACM Multimedia (ACM MM) Main Track Paper Accepted! Figure 1. Overview of the PEST framework. 🔎 TL;DR PEST introduces a novel framework for generalizable hateful meme moderation that leverages lightweight task-specific multimodal agents to steer powerful black-box vision-language models
Hugging Face Datasets2026 · Image
The Food API: Packaged Food & Beverage Products Dataset (Sample)The Food API: Packaged Food & Beverage Products Dataset Normalized packaged-food and beverage product data, keyed by barcode, for food-tech, nutrition apps and e-commerce. This repository is the free evaluation sample of The Food API: one record per product with the barcode, the full ingredient statement and its parsed tree, declared and precautionary allergens, the nutrition panel, package claims
Hugging Face Datasets2026 · Image
krotreaksmey/khmer-math-textbook🇰🇭 Khmer Math Textbook Line-Level OCR Dataset A large-scale, high-resolution dataset of line-level Khmer text and mathematical formulas extracted from official Cambodian Grade 9, 10, 11, and 12 mathematics textbooks. 📊 Dataset Summary Total Samples: 29,027 labeled line crops Coverage: Grade 9 (New): 6,847 line crops (math-G9-1 to math-G9-6847) Grades 10, 11, 12: 22,180 line crops Columns: Strictly
Hugging Face Datasets2026 · Image
The Food API: Packaged Food & Beverage Products Dataset (Sample)The Food API: Packaged Food & Beverage Products Dataset Normalized packaged-food and beverage product data, keyed by barcode, for food-tech, nutrition apps and e-commerce. This repository is the free evaluation sample of The Food API: one record per product with the barcode, the full ingredient statement and its parsed tree, declared and precautionary allergens, the nutrition panel, package claims
Hugging Face Datasets2026 · Image
Synthetic Text Images (English)Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise.
Hugging Face Datasets2026 · Image
AdvSpotAdvSpot AdvSpot is the first grounded adversarial OCR benchmark, introduced in the paper ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation. It targets visual text patterns that are readable to humans but challenging for multimodal large language models (MLLMs) to localize and recognize, and evaluates models with region-level annotations: bounding boxes,
Hugging Face Datasets2026 · Image
Malayalam OCR WordsMalayalam OCR Words A word-level Malayalam OCR dataset: cropped word images paired with their transcribed text label, split into train/validation/test sets. Dataset structure train.csv / val.csv / test.csv # tab-separated: <relative image path>\t<Malayalam word> train/ val/ test/ # image files referenced by the corresponding CSV Each CSV row maps one image file (path relative to its split folder)
Hugging Face Datasets2026 · Table · Parquet
mm_template_parquet(PCT 专利多模态语料)mm_template_parquet(PCT 专利多模态语料) PCT 专利公开文件的图文多模态语料,按 MNBVC 通用多模态模板 mm_template_mnbvc 的 BLOCK_SCHEMA 组织: 一块一行——一页一条 image-text-pair 块;无图但有 XML 的文档另出 text 块。 按目标体积分片:多个专利共用一个 parquet(目标 5GB/文件),不再是"一件专利一个文件"。 1000 件专利共 3 个分片;专利不跨文件,块ID 是实体内从 0 起的局部编号。 数据统计 项目 数值 专利(实体) 1,000 文档(源 ZIP) 20,084 页面图像 196,710 数据块 197,221(196,710 个 image-text-pair + 511 个 text) 数据大小 12,308,229,040 字节(约 11.5 GiB) 分片 3 个:4
Hugging Face Datasets2026 · Image
CADBench Extended Multimodal DatasetDataset Card Dataset Description CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually revi
Hugging Face Datasets2026 · Image
PDFA OCR Dataset - KREATIVE TIME BOXPDFA OCR Dataset Curated and Published by KREATIVE TIME BOX This dataset contains document page images along with their corresponding OCR layout bounding box annotations derived from PDFA document extraction pipelines. Dataset Overview Organization / Creator: KREATIVE TIME BOX Images: 27,499 PNG files (~8.8 GB) JSON Annotations: 6,989 JSON files (~52 MB) Image Format: PNG (RGB document page render
Hugging Face Datasets2026 · Image
Burmese Handwritten Sentence Dataset (BHSD)Burmese Handwritten Sentence Dataset (BHSD) BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR), error analysis, robustness testing, and research on low-resource scripts. The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers. This dataset would not have
Hugging Face Datasets2026 · Image
Synthetic Medical Document Recognition BenchmarkSynthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the d
Hugging Face Datasets2026 · Table · Parquet
Curated Danbooru 2026 Dataset (AVIF / Parquet)Curated Danbooru Streaming Dataset A large-scale, high-performance curated dataset of ~330,000 (330K) high-quality anime illustrations designed for training Diffusion Transformers (DiT), Latent Diffusion Models (LDM), and text-to-image generative models focused on the anime domain. This dataset is focused on specific curated characters and high-ranking artists using knowledge base lists (character
Hugging Face Datasets2026 · Image · gated
5CD-AI/VietHTR-LineWE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2026 · Image · gated
CwI-BenchCwI-Bench — Code-with-Image Bench 30 code-with-image tasks where models must write and run code against the image to reach the answer — pure visual inspection ("just look") is insufficient by design: answers require pixel-level precision (±px coordinates, exact counts, sub-degree angles, per-channel color recovery). Each task is a config; each config has three splits: split size source photos role
Hugging Face Datasets2026 · Image
Vasanthaleela/engineering-drawings-as1100Engineering Drawings AS1100 Compliance Dataset Dataset Description This dataset contains engineering drawings with various AS1100 (Australian Standard for Technical Drawing) compliance issues for training AI models to identify missing elements and non-compliance issues in technical drawings. Dataset Summary The Engineering Drawings AS1100 Compliance Dataset is designed to train and evaluate vision
Hugging Face Datasets2026 · Table · Parquet
zzzzineun/fundus-report-datasetFundus Report Dataset A curated benchmark dataset of ultra-widefield fundus photographs paired with ground-truth clinical reports, assembled from two public ophthalmology datasets. Each sample consists of a fundus image filename, a structured keyword label, and a free-text clinical report written by domain experts. The dataset is intended for evaluating medical vision-language models on fundus rep
Hugging Face Datasets2026 · Computed tomography
CancerVerse🩻 CancerVerse The first longitudinal, multimodal CT dataset spanning 13 malignant tumor types CancerVerse pairs whole-body abdominal/pelvic CT volumes with the radiologists' own free-text reports, follows patients across time, and is being released in stages — culminating in expert voxel-level tumor masks and a deep clinical & longitudinal annotation layer. It is built to power the next generation
Hugging Face Datasets2026 · dataset
DragOnDragOn: A Drag-Grounding Benchmark and Training Dataset for GUI Agents Drag-grounding dataset for GUI agents. Each example = one screenshot + one instruction + start/end bounding boxes for the drag. Four domains: text_highlight, sheet (cell selection), slide_resize (element resize/rotate/crop), slider. Dataset statistics Domain Train images Train tasks Eval (public) Dominant resolution Text Highli
Hugging Face Datasets2026 · Image
sarvamai/indic-ocr-benchSarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks,
Hugging Face Datasets2026 · Text
ud-synthetic/saudi-arabian-passportsDisclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Saudi Arabia The Synthetic Saudi Arabia Passports Dataset compiles more than 1,000 AI-generated passport images built for training OCR an
Hugging Face Datasets2026 · dataset · gated
Robo-Dopamine-GRM-DatasetRobo-Dopamine-GRM-Dataset A large-scale vision-language dataset for general robotic process reward modeling. Overview Robo-Dopamine studies general process reward modeling for robotic manipulation. The core GRM setting asks a vision-language model to compare robot states under a task instruction and estimate whether an AFTER state has made progress over a BEFORE state. This dataset provides image
Hugging Face Datasets2026 · Image · gated
5CD-AI/Viet-Handwriting-OCR-v2WE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2026 · Image
PureDocBenchMain Leaderboard 58 models · 3 matched tracks · 🏆 Search, filter & sort the leaderboard → The current evaluation covers 13 pipeline / multi-stage specialists, 19 end-to-end specialists, and 26 general-purpose VLMs. Each track contains 1,475 pages. The top 10 by the three-track mean, Avg₃, are shown below. Rank Model (release) Type Clean ↑ Digital ↑ Real ↑ Avg₃ ↑ 1 GLM-5.3-Flash (2026-08) General V