Constarium
← Search

Imaging · dataset · 2022

HuggingFaceM4/LocalizedNarratives

Listed in Hugging Face Datasets

Localized Narratives, a new form of multimodal image annotations connecting vision and language.

Description

We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the region they are describing. Since the voice and the mouse pointer are synchronized, we can localize every single word in the description.

This dense visual grounding takes the form of a mouse trace segment per word and is unique to our data. We annotated 849k images with Localized Narratives: the whole COCO, Flickr30k, and ADE20K datasets, and 671k images of Open Images, all of which we make publicly available.

Links

Documentation and papers

Catalogue records · 1

Topics

Stated by source
image
Inferred from text
Artificial intelligence 70% · Image 75%
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsHuggingFaceM4/LocalizedNarratives12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].anzsrc:group:4602enrichment · Hugging Facetaxonomy-embedding@1.0.0title+keywords+description (70%)
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:imagesource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:imageenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0