Hugging Face Datasets2026 · Table · Parquet
DeepASMR-NSpeechDeepASMR-NSpeech A Fine-Grained Benchmark for Non-Speech ASMR Generation 73,829 ten-second clips · 202.6 hours · 37 fine-grained actions 🎧 Interactive Demo · ⌘ Code · 📄 Paper: coming soon Dataset summary DeepASMR-NSpeech is a non-speech ASMR audio dataset with structured Subject-Verb-Object annotations and the SVO-AQA audio question-answering benchmark. All clips are 10 seconds long and cover 37 f
Hugging Face Datasets2026 · dataset
SoulX-Singer Eval DatasetSoulX-Singer-Eval This corpus contains 100 singing segments from 50 distinct individuals (25 Mandarin and 25 English speakers), with 2 segments provided per speaker. Additionally, 30 target samples are cross-selected from the Opencpop, M4Singer, and GTSinger, which are also in GMO-SVS. The annotation files are organized in JSONL format and categorized by text formulation (phoneme or word) and usag
Hugging Face Datasets2025 · Table · Parquet
TTS-Dataset-BatchedTest Version of humair025/TTS-Dataset-Batched TTS-Dataset-Batched Dataset Overview TTS-Dataset-Batched is a large-scale, multi-speaker English text-to-speech dataset optimized for efficient processing and training. The Original dataset contains 556,667 high-quality audio samples across 30 unique speakers, totaling over 1,024 hours of speech data. This is a batched version of a larger consolidated
Hugging Face Datasets2025 · Table · Parquet
jspaulsen/vctkVCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning. The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443. Reproducing This repository notably lacks a requirements.txt file. There's likely a
Hugging Face Datasets2025 · Text · gated
sander-wood/m4-rag🎵 M4-RAG: Million-scale Multilingual Music Metadata M4-RAG is a large-scale music-text dataset with 2.31 million music-text pairs, including 1.56 million audio-text pairs. It supports multimodal and multilingual music research, enabling tasks like text-to-music generation, music captioning, music information retrieval, and music classification. 🚀 🏆 Overview M4-RAG aggregates music metadata from di
Hugging Face Datasets2024 · Table · Parquet
ysdede/khanacademy-turkishKhan Academy Turkish Audio Dataset This dataset contains 78 hours of audio extracted from the Khan Academy Turkish YouTube channel. The data has been segmented into short clips, each with an average duration of 10.5 seconds. Accompanying this dataset, you will find a detailed video file tree that provides an overview of the source material. Dataset Creation Process:The audio was extracted from the
Hugging Face Datasets2024 · Table · Parquet
linagora/linto-dataset-audio-ar-tn-augmentedLinTO DataSet Audio for Arabic Tunisian Augmented A collection of Tunisian dialect audio and its annotations for STT task This is the augmented datasets used to train the Linto Tunisian dialect with code-switching STT linagora/linto-asr-ar-tn. Dataset Summary Dataset composition Sources Content Types Languages and Dialects Example use (python) License Citations Dataset Summary The LinTO DataSet Au
Hugging Face Datasets2024 · Table · Parquet
c張悦楷講古語音數據集 English 呢個係張悦楷講《三國演義》、《水滸傳》、《走進毛澤東的最後歲月》、《鹿鼎記》語音數據集。張悦楷係廣州最出名嘅講古佬 / 粵語説書藝人。佢從上世紀七十年代開始就喺廣東各個收音電台度講古,佢把聲係好多廣州人嘅共同回憶。本數據集收集嘅係佢最知名嘅四部作品。 數據集用途: TTS(語音合成)訓練集 ASR(語音識別)訓練集或測試集 各種語言學、文學研究 直接聽嚟欣賞藝術! TTS 效果演示:https://huggingface.co/spaces/laubonghaudoi/zoengjyutgaai_tts 説明 所有文本都根據 https://jyutping.org/blog/typo/ 同 https://jyutping.org/blog/particles/ 規範用字。 所有文本都使用全角標點,冇半角標點。 所有文本都用漢字轉寫,無阿拉伯數字無英
Hugging Face Datasets2022 · Table · Parquet · gated
GigaspeechDataset Card for Gigaspeech Dataset Description GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. Exam
Hugging Face Datasets2022 · dataset
VCTKThe CSTR VCTK Corpus includes speech data uttered by 110 English speakers with various accents.
Hugging Face Datasets2022 · dataset
LJ SpeechThis is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books in English. A transcription is provided for each clip. Clips vary in length from 1 to 10 seconds and have a total length of approximately 24 hours. Note that in order to limit the required storage for preparing this dataset, the audio is stored in the .wav for
Hugging Face Datasets2022 · Table · Parquet
MultiLingual LibriSpeechDataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - Engl