Data · dataset · 2026
Development of an Adaptive Text Summarization System Based on the T5 Transformer Model for Students with Dyslexia
Listed in Teesside University Research Data Repository
This dataset contains 110 Indonesian text samples developed to support research in automatic sign language translation, text simplification, and natural language processing.
Description
The data consist of original Indonesian texts accompanied by three levels of simplified target texts, namely Easy, Medium, and Hard. The Easy version uses simple vocabulary and short sentence structures to facilitate understanding, while the Medium version maintains more contextual information with moderate simplification.
The Hard version preserves the original meaning and sentence structure with minimal modifications. The dataset was manually compiled and annotated to simulate linguistic transformations commonly required in sign language translation systems, where complex sentence structures are simplified while maintaining semantic accuracy. Each record includes an identifier, original text, simplified target texts, source page information, and relevant keywords.
Read the rest (1 more)
This dataset is intended for applications such as text simplification, machine translation, sign language translation, accessibility technologies for deaf and hard-of-hearing communities, and the development of transformer-based language models including T5, mT5, BART, and IndoBART. The dataset is provided in Microsoft Excel (.xlsx) format and serves as a valuable resource for researchers working in Indonesian language processing, educational technology, and assistive communication systems.
Links
Where it is published
- DOI doi.org/10.17632/zv9d4x5vsp.1 ↗
DOI / persistent id · from researchdata tees ac uk
Catalogue records · 1
- OAI-PMH record data.mendeley.com/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Adata… ↗
metadata API · from researchdata tees ac uk
Topics
- From keywords
- Computer Science & AI · Earth & Environmental Science · Engineering · Humanities · Life Sciences · Natural language processing · Social Science
- Inferred from text
- Text 75%
Provenance · 1 source records, 13 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Teesside University Research Data Repository | oai:data.mendeley.com/zv9d4x5vsp.1 | 8 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].anzsrc:field:460208 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Natural Language Processing'] |
| concepts[field].local:field:computer-science-ai | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:engineering | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:humanities | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:social-science | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[modality].local:modality:text | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/description |
| license | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/rights |
| publication_date | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| title | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/title |