Data · dataset · 2026
Darbest Dataset: Universal Dependencies Treebank for Standard Sorani Kurdish
Listed in Teesside University Research Data Repository
The Darbest dataset is a Universal Dependencies (UD) treebank dataset for Standard Sorani Kurdish written in the Perso-Arabic script.
Description
It contains 69,000 annotated sentences and 1,205,855 tokens collected from nine textual domains. The corpus was collected from seven Kurdish online news websites and supplemented with texts from published books.
Before preprocessing, the collected corpus contained 1,250,275 words from 5,627 web pages together with book-based texts and was preprocessed using a Python-based pipeline involving text cleaning, Unicode and punctuation normalization, sentence segmentation, and tokenization. The dataset was developed to provide a large-scale syntactically and morphologically annotated resource for Standard Sorani Kurdish. A separate 100-sentence gold-standard set was manually annotated according to the Universal Dependencies v2 guidelines.
Read the rest (2 more)
Sorani Kurdish linguistic experts supported the selection of sentences representing diverse and linguistically complex structures and reviewed LLM-generated annotations for errors. The gold-standard set was used to construct few-shot prompts for annotating the remaining corpus. The resulting annotations were represented in the standard CoNLL-U format and validated using the official Universal Dependencies validation tool, followed by manual correction and quality review.
The released treebank is divided into training, development, and test sets and includes lemmas, Universal Part-of-Speech (UPOS) tags, morphological features, syntactic heads, and dependency relations. The dataset can be used to train, evaluate, and benchmark NLP models for part-of-speech tagging, lemmatization, morphological analysis, dependency parsing, and related computational linguistics tasks. The accompanying repository also contains the separate 100-sentence gold-standard set, plain-text corpus splits, README documentation, and a dataset statistics spreadsheet.
Links
Where it is published
- DOI doi.org/10.17632/msy32gzj9j.2 ↗
DOI / persistent id · from researchdata tees ac uk
Catalogue records · 1
- OAI-PMH record data.mendeley.com/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Adata… ↗
metadata API · from researchdata tees ac uk
Topics
- From keywords
- Computational linguistics · Computer Science & AI · Earth & Environmental Science · Engineering · Humanities · Life Sciences · Linguistics · Natural language processing · Social Science
- Inferred from text
- Text 75%
Provenance · 1 source records, 15 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Teesside University Research Data Repository | oai:data.mendeley.com/msy32gzj9j.2 | 4 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].anzsrc:field:460208 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Natural Language Processing'] |
| concepts[field].anzsrc:field:470403 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Computational Linguistics'] |
| concepts[field].anzsrc:group:4704 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Linguistics'] |
| concepts[field].local:field:computer-science-ai | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:engineering | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:humanities | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:social-science | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[modality].local:modality:text | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/description |
| license | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/rights |
| publication_date | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| title | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/title |