Constarium
← Search

Imaging · dataset · 2023

COCO-style geographically unbiased image dataset for computer vision applications

Listed in DataSuds

There are already a lot of datasets linked to computer vision tasks (Imagenet, MS COCO, Pascal VOC, OpenImages, and numerous others), but they all suffer from important bias.

Description

One bias of significance for us is the data origin: most datasets are composed of data coming from developed countries. Facing this situation, and the need of data with local context in developing countries, we try here to adapt common data generation process to inclusive data, meaning data drawn from locations and cultural context that are unseen or poorly represented.

We chose to replicate MS COCO's data generation process, as it is well documented and easy to implement. Data was collected from January to April 2022 through Flickr platform. This dataset contains the results of our data collection process, as follows : 23 text files containing comma separated URLs for each of the 23 geographic zones identified in the UN M49 norm.

Read the rest (3 more)

These text files are named according to the names of the geographic zones they cover. Annotations for 400 images per geographic zones. Those annotations are COCO-style, and inform on the presence or absence of 91 categories of objects or concepts on the images.

They are shared in a JSON format. Licenses for the 400 annotations per geographic zones, based on the original licenses of the data and specified per image. Those licenses are shared under CSV format.

A document explaining the objectives and methodology underlying the data collection, also describing the different components of the dataset.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Text 75%
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
DataSudsdoi:10.23708/N2UY4C8 d agoJSON v1
FieldAssertionExtractorEvidence
concepts[field].anzsrc:field:460304mapping · dataverse ird frvocabulary-mapper@1.0.0keywords['Computer Vision']
concepts[field].dataverse_subject:computer-and-information-sciencesource · dataverse ird frconnector:dataverse_ird_fr@1.0.0/subjects
concepts[field].local:field:computer-science-aimapping · dataverse ird frconnector:dataverse_ird_fr@1.0.0/subjects
concepts[modality].local:modality:imagemapping · dataverse ird frvocabulary-mapper@1.0.0keywords['Images']
concepts[modality].local:modality:textenrichment · dataverse ird frkeyword-concept-rules@1.0.0title+description (75%)
created_datesource · dataverse ird frconnector:dataverse_ird_fr@1.0.0
descriptionsource · dataverse ird frconnector:dataverse_ird_fr@1.0.0/description
publication_datesource · dataverse ird frconnector:dataverse_ird_fr@1.0.0
titlesource · dataverse ird frconnector:dataverse_ird_fr@1.0.0/name
updated_datesource · dataverse ird frconnector:dataverse_ird_fr@1.0.0
version_labelsource · dataverse ird frconnector:dataverse_ird_fr@1.0.0