Constarium
← Search

Text · dataset · 2025

ChildMandarin

Listed in Hugging Face Datasets

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5 Introduction ChildMandarin is a comprehensive, open-source Mandarin Chinese speech dataset specifically designed for research on young children aged 3 to 5.

Description

This dataset addresses the critical lack of publicly available resources for this age group, enabling advancements in automatic speech recognition (ASR), speaker verification (SV), and other related fields.

The dataset is released… See the full description on the dataset page: huggingface.co/datasets/BAAI/ChildMandarin.

Links

Get the data

Documentation and papers

Catalogue records · 1

Topics

Inferred from text
Audio 65%
Provenance · 1 source records, 12 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsBAAI/ChildMandarin12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:audiosource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:audioenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (65%)
concepts[task].hf_task:automatic-speech-recognitionsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0