Constarium
← Search

Data · dataset · 2025

UGAkan-ImpairedSpeechData: A Dataset of Impaired Speech in the Akan Language

Listed in Teesside University Research Data Repository

The UGAkan-ImpairedSpeechData is a speech dataset from indigenous speakers of Akan with different forms of speech impairments.

Description

It contains audio descriptions of culturally relevant images. The dataset comprises 14,312 audio files and corresponding transcriptions equivalent to 50.01 hours.

Recordings were done in different environments including Outdoor (7,706 audio files), Other (3,075), Indoor (2,254), Studio (982), and Car (295). The dataset is also categorized by aetiology and gender. Male speakers contributed 6,754 files equivalent to 19.02 hours, with the highest representation from individuals with Cerebral Palsy (2,881 files, 8.84 hours), followed by Stammering, Cleft, and Stroke.

Read the rest (2 more)

Female speakers contributed 7,558 files equivalent to 30.99 hours, with most recordings coming from individuals with Cerebral Palsy (4,835 files, 15.66 hours) and Stammering (2,574 files, 13.88 hours). Stroke data was recorded only from male speakers, while Cleft speech samples were collected from both genders, with a higher volume from males. In terms of duration, the audio files vary in length.

The average audio length is 12.46 seconds, with a standard deviation of 7.71 seconds, indicating moderate variability. The majority of audio files range from 6.59s to 16.00s, suggesting a right-skewed distribution. The maximum duration is 60.08s, which exceeds the upper bound of the interquartile range and is likely an outlier.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Audio 75% · Image 65%
Provenance · 1 source records, 16 field assertions
SourceKeyLast seenRaw
Teesside University Research Data Repositoryoai:data.mendeley.com/vc84vdw8tb.44 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].anzsrc:field:460208mapping · researchdata tees ac ukvocabulary-mapper@1.0.0keywords['Natural Language Processing']
concepts[field].anzsrc:field:460212mapping · researchdata tees ac ukvocabulary-mapper@1.0.0keywords['Speech Recognition']
concepts[field].anzsrc:group:4704mapping · researchdata tees ac ukvocabulary-mapper@1.0.0keywords['Linguistics']
concepts[field].local:field:computer-science-aimapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:earth-environmentalmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:engineeringmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:humanitiesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:life-sciencesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:social-sciencemapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[modality].local:modality:audioenrichment · researchdata tees ac ukkeyword-concept-rules@1.0.0title+description (75%)
concepts[modality].local:modality:imageenrichment · researchdata tees ac ukkeyword-concept-rules@1.0.0title+description (65%)
descriptionsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/description
licensesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/rights
publication_datesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
titlesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/title