Constarium
← Search

Data · dataset · 2026

ShadhuCholito-BN

Listed in Teesside University Research Data Repository

**Dataset Description** This dataset contains parallel Bangla and English sentence pairs collected and processed for language translation and linguistic research.

Description

The Bangla corpus includes sentences written in both **Cholito Bangla** (colloquial form) and **Sadhu Bangla** (classical/formal form), extracted from Bangla literary sources. Each sentence is labeled according to its linguistic style and paired with an English translation.

The dataset consists of the following fields:

Read the rest (1 more)
  • **sentence**: The sentence text (Bangla or English, depending on the dataset version).
  • **label**: Linguistic style label (`cholito` or `sadhu`).
  • **source**: Source category of the text (`Book`). Data preprocessing included sentence segmentation using Bangla punctuation markers, removal of noisy OCR artifacts, duplicate elimination, and dataset validation. The English version was generated through machine translation and aligned with the original Bangla sentences. The dataset is intended for research and development in:
  • Machine Translation (Bangla–English)
  • Natural Language Processing (NLP)
  • Text Classification
  • Style Transfer
  • Bangla Linguistic Analysis
  • Large Language Model (LLM) Training and Evaluation The dataset was randomly shuffled and split into training and testing subsets using a 70:30 ratio while preserving sentence alignment between the Bangla and English versions.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Optical character recognition 65% · Text 75%
Provenance · 1 source records, 14 field assertions
SourceKeyLast seenRaw
Teesside University Research Data Repositoryoai:data.mendeley.com/fyrcfp598s.18 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].anzsrc:field:460208mapping · researchdata tees ac ukvocabulary-mapper@1.0.0keywords['Natural Language Processing']
concepts[field].local:field:computer-science-aimapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:earth-environmentalmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:engineeringmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:humanitiesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:life-sciencesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:social-sciencemapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[method].local:method:ocrenrichment · researchdata tees ac ukkeyword-concept-rules@1.0.0title+description (65%)
concepts[modality].local:modality:textenrichment · researchdata tees ac ukkeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/description
licensesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/rights
publication_datesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
titlesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/title