Constarium
← Search

Data · dataset · 2026

Development of an Adaptive Text Summarization System Based on the T5 Transformer Model for Students with Dyslexia

Listed in Teesside University Research Data Repository

This dataset contains 110 Indonesian text samples developed to support research in automatic sign language translation, text simplification, and natural language processing.

Description

The data consist of original Indonesian texts accompanied by three levels of simplified target texts, namely Easy, Medium, and Hard. The Easy version uses simple vocabulary and short sentence structures to facilitate understanding, while the Medium version maintains more contextual information with moderate simplification.

The Hard version preserves the original meaning and sentence structure with minimal modifications. The dataset was manually compiled and annotated to simulate linguistic transformations commonly required in sign language translation systems, where complex sentence structures are simplified while maintaining semantic accuracy. Each record includes an identifier, original text, simplified target texts, source page information, and relevant keywords.

Read the rest (1 more)

This dataset is intended for applications such as text simplification, machine translation, sign language translation, accessibility technologies for deaf and hard-of-hearing communities, and the development of transformer-based language models including T5, mT5, BART, and IndoBART. The dataset is provided in Microsoft Excel (.xlsx) format and serves as a valuable resource for researchers working in Indonesian language processing, educational technology, and assistive communication systems.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Text 75%
Provenance · 1 source records, 13 field assertions
SourceKeyLast seenRaw
Teesside University Research Data Repositoryoai:data.mendeley.com/zv9d4x5vsp.18 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].anzsrc:field:460208mapping · researchdata tees ac ukvocabulary-mapper@1.0.0keywords['Natural Language Processing']
concepts[field].local:field:computer-science-aimapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:earth-environmentalmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:engineeringmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:humanitiesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:life-sciencesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:social-sciencemapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[modality].local:modality:textenrichment · researchdata tees ac ukkeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/description
licensesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/rights
publication_datesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
titlesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/title