Constarium
← Search

Data · dataset · 2026

Multilingual AI-Based Dyslexia Detection

Listed in Teesside University Research Data Repository

This dataset contains 852 character-level handwriting images collected from 100 unique children aged 6–10 years in school settings in Narowal, Punjab, Pakistan, for research on handwriting-based dyslexia screening.

Description

The dataset is balanced into two classes: 426 images from the dyslexic/YES group and 426 images from the non-dyslexic/NO group. Participants completed handwriting tasks covering the English alphabet, Urdu alphabet, and numerals 0–9 using standard paper and pens.

Handwritten samples were digitized at character level using an Apple iPhone 11 (12 MP) camera. Images were visually inspected for sharpness, visibility, completeness, and correct character boundaries. Severely blurred, incomplete, or incorrectly cropped images were re-captured or excluded.

Read the rest (3 more)

Urdu handwriting images were retained without OCR-based correction or linguistic normalization to preserve visually relevant characteristics such as dots, curves, loops, stroke formation, spacing, and character morphology. The deposited collection includes the accepted handwriting images together with a metadata CSV and README documentation. The metadata records image identifiers, original filenames, class labels, image dimensions, camera information where available, file format, file size, quality-control status, preprocessing information, and integrity checksums.

Participant identifiers are intended to remain anonymized. The dataset can support research in computer vision, deep learning, transfer learning, image classification, handwriting analysis, educational AI, low-resource language processing, and bilingual or multi-script dyslexia-screening methods. It may also be used to compare preprocessing, feature-learning, augmentation, validation, and classification strategies.

This dataset is associated with the published research article: Kashif, M., Haider, Z. M., Muneer, I., Kumar, D., Tahir, T., & Shafi, J. (2026). “Optimizing Deep Neural Models for Early Dyslexia Detection Using Novel Bilingual Handwritten Dataset.” Engineering Reports, 8, e70867. doi.org/10.1002/eng2.70867. Because the dataset concerns handwriting collected from children, reuse should comply with the applicable ethical approval, consent conditions, privacy safeguards, and repository access conditions.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Image 75%
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Teesside University Research Data Repositoryoai:data.mendeley.com/b37s8b94nv.14 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:earth-environmentalmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:engineeringmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:humanitiesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:life-sciencesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:social-sciencemapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[modality].local:modality:imageenrichment · researchdata tees ac ukkeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/description
licensesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/rights
publication_datesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
titlesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/title