Data · dataset · 2026
ShadhuCholito-BN
Listed in Teesside University Research Data Repository
**Dataset Description** This dataset contains parallel Bangla and English sentence pairs collected and processed for language translation and linguistic research.
Description
The Bangla corpus includes sentences written in both **Cholito Bangla** (colloquial form) and **Sadhu Bangla** (classical/formal form), extracted from Bangla literary sources. Each sentence is labeled according to its linguistic style and paired with an English translation.
The dataset consists of the following fields:
Read the rest (1 more)
- **sentence**: The sentence text (Bangla or English, depending on the dataset version).
- **label**: Linguistic style label (`cholito` or `sadhu`).
- **source**: Source category of the text (`Book`). Data preprocessing included sentence segmentation using Bangla punctuation markers, removal of noisy OCR artifacts, duplicate elimination, and dataset validation. The English version was generated through machine translation and aligned with the original Bangla sentences. The dataset is intended for research and development in:
- Machine Translation (Bangla–English)
- Natural Language Processing (NLP)
- Text Classification
- Style Transfer
- Bangla Linguistic Analysis
- Large Language Model (LLM) Training and Evaluation The dataset was randomly shuffled and split into training and testing subsets using a 70:30 ratio while preserving sentence alignment between the Bangla and English versions.
Links
Where it is published
- DOI doi.org/10.17632/fyrcfp598s.1 ↗
DOI / persistent id · from researchdata tees ac uk
Catalogue records · 1
- OAI-PMH record data.mendeley.com/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Adata… ↗
metadata API · from researchdata tees ac uk
Topics
- From keywords
- Computer Science & AI · Earth & Environmental Science · Engineering · Humanities · Life Sciences · Natural language processing · Social Science
- Inferred from text
- Optical character recognition 65% · Text 75%
Provenance · 1 source records, 14 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Teesside University Research Data Repository | oai:data.mendeley.com/fyrcfp598s.1 | 8 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].anzsrc:field:460208 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Natural Language Processing'] |
| concepts[field].local:field:computer-science-ai | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:engineering | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:humanities | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:social-science | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[method].local:method:ocr | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (65%) |
| concepts[modality].local:modality:text | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/description |
| license | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/rights |
| publication_date | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| title | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/title |