Data · dataset · 2023
HugoLaurencon/IIIT-5K
Listed in Hugging Face Datasets
The IIIT 5K-Word dataset is harvested from Google image search.
Description
Query words like billboards, signboard, house numbers, house name plates, movie posters were used to collect images. The dataset contains 5000 cropped word images from Scene Texts and born-digital images.
The dataset is divided into train and test parts. This dataset can be used for large lexicon cropped word recognition. We also provide a lexicon of more than 0.5 million dictionary words with this dataset.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/HugoLaurencon/IIIT-5K ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/HugoLaurencon/IIIT-5K ↗
metadata API · from Hugging Face
Topics
- From keywords
- Computer Science & AI
- Inferred from text
- Image 75%
Provenance · 1 source records, 8 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | HugoLaurencon/IIIT-5K | 12 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].local:modality:image | enrichment · Hugging Face | keyword-concept-rules@1.0.0 | title+description (75%) |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |