Constarium
← Search

Data · dataset · 2022

guoqiang/cuge

Listed in Hugging Face Datasets

Dataset

Description

Summary

The Common Voice dataset consists of a unique MP3 and corresponding text file. Many of the 9,283 recorded hours in the dataset also include demographic metadata like age, sex, and accent that can help train the accuracy of speech recognition engines. The dataset currently consists of 7,335 validated hours in 60 languages, but were always adding more voices and languages.

Read the rest (1 more)

Take a look at our Languages page to request a language or start contributing. Supported Tasks and… See the full description on the dataset page: huggingface.co/datasets/guoqiang/cuge.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Audio 65% · Text 75%
Provenance · 1 source records, 9 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsguoqiang/cuge12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:audioenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (65%)
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0