Constarium
← Search

Data · dataset · 2022

Common Language

Listed in Hugging Face Datasets

This dataset is composed of speech recordings from languages that were carefully selected from the CommonVoice database.

Description

The total duration of audio recordings is 45.1 hours (i.e., 1 hour of material for each language). The dataset has been extracted from CommonVoice to train language-id systems.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
audio classification
Inferred from text
Audio 75%
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsspeechbrain/common_language12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:audioenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[task].hf_task:audio-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0