Imaging · dataset · 2026
PDFA OCR Dataset - KREATIVE TIME BOX
Listed in Hugging Face Datasets
Description
PDFA OCR Dataset Curated and Published by KREATIVE TIME BOX This dataset contains document page images along with their corresponding OCR layout bounding box annotations derived from PDFA document extraction pipelines. Dataset Overview Organization / Creator: KREATIVE TIME BOX Images: 27,499 PNG files (~8.8 GB) JSON Annotations: 6,989 JSON files (~52 MB) Image Format: PNG (RGB document page renders) Annotation Format: JSON with text lines, normalized bounding… See the full description on the dataset page: huggingface.co/datasets/infokreativetimebox/ktb-ocr-dataset.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/infokreativetimebox/ktb-ocr-dataset ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/infokreativetimebox/ktb-ocr-dataset ↗
metadata API · from Hugging Face
Topics
- Stated by source
- image · image to text · object detection · visual question answering
- From keywords
- Computer Science & AI · Optical character recognition
Provenance · 1 source records, 15 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | infokreativetimebox/ktb-ocr-dataset | 6 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[method].local:method:ocr | mapping · Hugging Face | vocabulary-mapper@1.0.0 | keywords['ocr'] |
| concepts[modality].hf_modality:image | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].local:modality:image | enrichment · Hugging Face | keyword-concept-rules@1.0.0 | title+description (75%) |
| concepts[modality].local:modality:text | enrichment · Hugging Face | keyword-concept-rules@1.0.0 | title+description (75%) |
| concepts[task].hf_task:image-to-text | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| concepts[task].hf_task:object-detection | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| concepts[task].hf_task:visual-question-answering | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| license | source · Hugging Face | connector:huggingface@1.0.0 | /tags[license:*] |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |