Constarium
← Search

Data · dataset · 2022

Modelling word learning and recognition using visually grounded speech

Listed in DANS Data Station Social Sciences and Humanities

A set of recorded isolated nouns, verbs and image annotations used for testing the word recognition performance of our speech2image model.

Description

We trained a word recognition model on a set of images and utterances. The model should learn to recognise words without ever having seen written transcripts.

The word recognition performance is measured as the number of retrieved images out of 10 displaying the correct visual referent. We furthermore collected new ground truth object and action annotations for the Flickr8k test images for this purposes. This consists of 1000 images, all annotated for the presence of the 50 actions and objects corresponding to the test verbs and nouns.

Read the rest (2 more)

In order to test the word recognition performance we took the 50 most common nouns and 50 most common verbs in the training data, confirmed that there were at least 10 images in our test image data that displayed these actions and objects. These nouns and verbs where recorded in singular and plural form (nouns) and in root, third person and progressive form (verbs). We furthermore annotated 1000 images from the Flickr8k test set for the presence of these nouns and verbs.

These annotations are included in .CSV format

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Audio 65% · Image 75%
Provenance · 1 source records, 14 field assertions
SourceKeyLast seenRaw
DANS Data Station Social Sciences and Humanitiesdoi:10.17026/DANS-22N-XH474 d agoJSON v1
FieldAssertionExtractorEvidence
concepts[field].anzsrc:field:460212mapping · ssh datastations nlvocabulary-mapper@1.0.0keywords['speech recognition']
concepts[field].dataverse_subject:arts-and-humanitiessource · ssh datastations nlconnector:ssh_datastations_nl@1.0.0/subjects
concepts[field].dataverse_subject:computer-and-information-sciencesource · ssh datastations nlconnector:ssh_datastations_nl@1.0.0/subjects
concepts[field].local:field:computer-science-aimapping · ssh datastations nlconnector:ssh_datastations_nl@1.0.0/subjects
concepts[field].local:field:humanitiesmapping · ssh datastations nlconnector:ssh_datastations_nl@1.0.0/subjects
concepts[field].local:field:social-sciencemapping · ssh datastations nlconnector:ssh_datastations_nl@1.0.0/subjects
concepts[modality].local:modality:audioenrichment · ssh datastations nlkeyword-concept-rules@1.0.0title+description (65%)
concepts[modality].local:modality:imageenrichment · ssh datastations nlkeyword-concept-rules@1.0.0title+description (75%)
created_datesource · ssh datastations nlconnector:ssh_datastations_nl@1.0.0
descriptionsource · ssh datastations nlconnector:ssh_datastations_nl@1.0.0/description
publication_datesource · ssh datastations nlconnector:ssh_datastations_nl@1.0.0
titlesource · ssh datastations nlconnector:ssh_datastations_nl@1.0.0/name
updated_datesource · ssh datastations nlconnector:ssh_datastations_nl@1.0.0
version_labelsource · ssh datastations nlconnector:ssh_datastations_nl@1.0.0