Constarium
← Search

Data · dataset · 2023

HugoLaurencon/IIIT-5K

Listed in Hugging Face Datasets

The IIIT 5K-Word dataset is harvested from Google image search.

Description

Query words like billboards, signboard, house numbers, house name plates, movie posters were used to collect images. The dataset contains 5000 cropped word images from Scene Texts and born-digital images.

The dataset is divided into train and test parts. This dataset can be used for large lexicon cropped word recognition. We also provide a lexicon of more than 0.5 million dictionary words with this dataset.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Image 75%
Provenance · 1 source records, 8 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsHugoLaurencon/IIIT-5K12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:imageenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0