Constarium
← Search

Data · dataset · 2026

BanglaNER-Corp: A Novel Reference Corpus for Bangla Named Entity Recognition

Listed in Teesside University Research Data Repository

This repository contains a Bangla Named Entity Recognition (NER) dataset along with the source code used for data preprocessing.

Description

The dataset is a hybrid Bangla Named Entity Recognition (NER) corpus designed to support research on low-resource natural language processing. It combines translated and native Bangla text to provide a diverse and high-quality benchmark for training and evaluating NER models.

Approximately 70% of the corpus was created by translating the English CoNLL-2003 dataset into Bangla using neural machine translation models (BanglaT5, NLLB, and Google Translate), followed by manual verification and correction to preserve semantic meaning and linguistic fluency. The original BIO entity annotations for Person (PER), Organization (ORG), Location (LOC), and Miscellaneous (MISC) were retained and carefully realigned with the translated text.

Read the rest (3 more)

The remaining 30% consists of manually collected and annotated Bangla news sentences spanning multiple domains, including politics, sports, international affairs, culture, weather, mythology, religion, science, and geography. The complete corpus contains 6,162 annotated sentences and 69,193 tokens, with a vocabulary of 14,349 unique tokens. Entity annotations include 4,609 PER, 2,904 ORG, 4,678 LOC, and 3,172 MISC tokens, alongside 53,830 non-entity (O) tokens.

All annotations follow the BIO tagging scheme and are provided in CoNLL format, making the dataset directly compatible with modern NLP frameworks such as Hugging Face Transformers, PyTorch, TensorFlow, and spaCy. In addition to the dataset, this repository includes the complete source code for data preprocessing. This dataset is intended for training, evaluation, benchmarking, and comparative research on Bangla NER systems, particularly transformer-based and hybrid sequence-labeling models.

It also serves as a valuable resource for downstream Bangla NLP applications, including information extraction, question answering, and knowledge graph construction.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Text 75%
Provenance · 1 source records, 14 field assertions
SourceKeyLast seenRaw
Teesside University Research Data Repositoryoai:data.mendeley.com/dy4nhv5tj2.27 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].anzsrc:field:460208mapping · researchdata tees ac ukvocabulary-mapper@1.0.0keywords['Natural Language Processing']
concepts[field].anzsrc:group:4602mapping · researchdata tees ac ukvocabulary-mapper@1.0.0keywords['Artificial Intelligence']
concepts[field].local:field:computer-science-aimapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:earth-environmentalmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:engineeringmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:humanitiesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:life-sciencesmapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[field].local:field:social-sciencemapping · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
concepts[modality].local:modality:textenrichment · researchdata tees ac ukkeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/description
licensesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/rights
publication_datesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0
titlesource · researchdata tees ac ukconnector:researchdata_tees_ac_uk@1.0.0/metadata/dc/title