Data · dataset · 2026
A Synchronized Chittagonian Speech Corpus for Dialect ASR and Speech-to-Speech Translation
Listed in Teesside University Research Data Repository
This dataset presents a synchronized Chittagonian speech corpus developed for low-resource speech processing and dialect translation.
Description
The corpus contains 8,004 Chittagonian speech samples paired with corresponding Chittagonian transcripts and Standard Bangla translations. The textual component of the dataset is taken from the publicly available Kothon dataset, a Chittagonian–Standard Bangla parallel text resource.
In this work, the original text pairs were extended with newly collected speech recordings from native Chittagonian speakers, creating a synchronized speech-text corpus for automatic speech recognition (ASR), speech-to-speech translation, machine translation, dialect normalization, and text-to-speech (TTS) . The dataset is organized into predefined training, validation, and test splits containing 6,403, 801, and 800 samples, respectively.
Read the rest (1 more)
Each audio recording is provided in WAV format. This resource aims to support reproducible research on Chittagonian, an under-resourced regional variety of Bangla, by enabling the development and evaluation of machine learning models for dialect-aware speech and language technologies.
Links
Where it is published
- DOI doi.org/10.17632/n9xttk45df.1 ↗
DOI / persistent id · from researchdata tees ac uk
Catalogue records · 1
- OAI-PMH record data.mendeley.com/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Adata… ↗
metadata API · from researchdata tees ac uk
Topics
Provenance · 1 source records, 15 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Teesside University Research Data Repository | oai:data.mendeley.com/n9xttk45df.1 | 4 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].anzsrc:field:460208 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Natural Language Processing'] |
| concepts[field].anzsrc:field:460212 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Speech Recognition'] |
| concepts[field].local:field:computer-science-ai | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:engineering | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:humanities | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:social-science | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[modality].local:modality:audio | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| concepts[modality].local:modality:text | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/description |
| license | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/rights |
| publication_date | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| title | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/title |