Constarium
← Search

Imaging · dataset · 2026

sarvamai/indic-ocr-bench

Listed in Hugging Face Datasets

Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor.

Description

Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material.

The benchmark covers 23… See the full description on the dataset page: huggingface.co/datasets/sarvamai/indic-ocr-bench.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
image · image to text · text · text generation
Provenance · 1 source records, 13 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetssarvamai/indic-ocr-bench10 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[method].local:method:ocrmapping · Hugging Facevocabulary-mapper@1.0.0keywords['ocr']
concepts[modality].hf_modality:imagesource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:image-to-textsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0