{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:document-question-answering",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Image
DuckerMaster/Thai-Synth-ReceiptsThai-Synth-Receipts Thai-Synth-Receipts is a large-scale, highly robust synthetic dataset of Thai commercial documents designed specifically for training and evaluating state-of-the-art Document AI and Optical Character Recognition (OCR) models. The dataset consists of 14,976 high-resolution document images (Receipts, Tax Invoices, Thermal Slips, and Quotations) across three distinct degradation v
Hugging Face Datasets2026 · Image
Synthetic Medical Document Recognition BenchmarkSynthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the d
Hugging Face Datasets2026 · Image · gated
PopResumePopResume Population-Representative Resume Dataset for Causal Fairness Evaluation of LLM/VLM Resume Screeners 🌐 Project page · 📄 arXiv · 💻 Code · EMNLP 2026 Main (Oral) PopResume is a synthetic, population-representative resume dataset that preserves the natural statistical relationships among demographic attributes, demographic proxies, and job-relevant qualifications found in U.S. population sta
Hugging Face Datasets2026 · Image
PureDocBenchMain Leaderboard 58 models · 3 matched tracks · 🏆 Search, filter & sort the leaderboard → The current evaluation covers 13 pipeline / multi-stage specialists, 19 end-to-end specialists, and 26 general-purpose VLMs. Each track contains 1,475 pages. The top 10 by the three-track mean, Avg₃, are shown below. Rank Model (release) Type Clean ↑ Digital ↑ Real ↑ Avg₃ ↑ 1 GLM-5.3-Flash (2026-08) General V
Hugging Face Datasets2026 · Image
yuheng-chen/Unified_Document_Understanding_DatasetDataset Card for Read-Parsing-Describe: Unified Scientific Document Understanding Read-Parsing-Describe (Unified Scientific Document Understanding) is a pioneering multimodal benchmark designed to train and evaluate models on the complex structures of scientific documents. Unlike traditional document datasets that treat visual elements merely as isolated layout blocks, RPD transforms document pars
Hugging Face Datasets2023 · Text
niftyThe News-Informed Financial Trend Yield (NIFTY) Dataset. The News-Informed Financial Trend Yield (NIFTY) Dataset. Details of the dataset, including data procurement and filtering can be found in the paper here: https://arxiv.org/abs/2405.09747. For the NIFTY-RL LLM alignment dataset please use nifty-rl. 📋 Table of Contents 🧩 NIFTY Dataset 📋 Table of Contents 📖 Usage Downloading the dataset Dataset