Data · dataset · 2026
AranjiyyaCorpus: A span-annotated Arabic news dataset of anglicised style, calques, and borrowings across six domains
Listed in Teesside University Research Data Repository
Description
The Arabic neologism Aranjiyya (عَرَنْجِيَّة) is a portmanteau of ʿArabiyya (Arabic) and Inkliziyya / Faranjiyya (English/foreign), used by contemporary Arabic editors and stylists to describe Arabic prose that retains Arabic vocabulary while importing the syntactic, stylistic, semantic, or lexical structures of English. The phenomenon is pervasive in translated news, press releases, technical writing, and digital media, and is a recurrent target of prescriptive Arabic-style guides.
Despite its prominence, Aranjiyya has had almost no presence in computational Arabic resources: existing treebanks and error corpora target orthographic, morphological, or syntactic well-formedness but do not isolate contact-induced patterns whose surface forms are grammatical but whose underlying templates are English. AranjiyyaCorpus was constructed to fill this gap, with three motivating use cases: training a span-level Aranjiyya detector for editors and translation post-editors; producing evaluation data for whether large language models actually generate idiomatic Arabic; and supporting linguistic study of contact-induced change in modern Arabic, with sufficient category granularity to distinguish syntactic, stylistic, and semantic phenomena and to track them across genres.
Links
Where it is published
- DOI doi.org/10.17632/rh8gf885hz.1 ↗
DOI / persistent id · from researchdata tees ac uk
Catalogue records · 1
- OAI-PMH record data.mendeley.com/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Adata… ↗
metadata API · from researchdata tees ac uk
Topics
Provenance · 1 source records, 13 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Teesside University Research Data Repository | oai:data.mendeley.com/rh8gf885hz.1 | 7 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].anzsrc:field:460208 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Natural Language Processing'] |
| concepts[field].anzsrc:field:470403 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Computational Linguistics'] |
| concepts[field].local:field:computer-science-ai | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:engineering | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:humanities | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:social-science | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| description | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/description |
| license | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/rights |
| publication_date | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| title | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/title |