Constarium
← Search

Table · dataset · 2023

SMC Malayalam Speech Corpus

Listed in Hugging Face Datasets

Description

Dataset Card for [msc] Dataset Summary 1541 speech samples 75 speech contributors 1:38:16 hours of speech 482 unique sentences 1400 unique words 553 unique syllables 48 unique phonemes For more detailed analysis see the python notebook provided here Supported Tasks and Leaderboards Automatic Speech Recognition system development, gender and age identification of speakers Languages Malayalam Dataset Structure file_name… See the full description on the dataset page: huggingface.co/datasets/smcproject/MSC.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Artificial intelligence 72% · Audio 65%
Provenance · 1 source records, 13 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetssmcproject/MSC11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].anzsrc:group:4602enrichment · Hugging Facetaxonomy-embedding@1.1.0title+keywords+description (72%)
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:audiosource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:audioenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (65%)
concepts[task].hf_task:automatic-speech-recognitionsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0