Data · dataset · 2026
BanglaNER-Corp: A Novel Reference Corpus for Bangla Named Entity Recognition
Listed in Teesside University Research Data Repository
This repository contains a Bangla Named Entity Recognition (NER) dataset along with the source code used for data preprocessing.
Description
The dataset is a hybrid Bangla Named Entity Recognition (NER) corpus designed to support research on low-resource natural language processing. It combines translated and native Bangla text to provide a diverse and high-quality benchmark for training and evaluating NER models.
Approximately 70% of the corpus was created by translating the English CoNLL-2003 dataset into Bangla using neural machine translation models (BanglaT5, NLLB, and Google Translate), followed by manual verification and correction to preserve semantic meaning and linguistic fluency. The original BIO entity annotations for Person (PER), Organization (ORG), Location (LOC), and Miscellaneous (MISC) were retained and carefully realigned with the translated text.
Read the rest (3 more)
The remaining 30% consists of manually collected and annotated Bangla news sentences spanning multiple domains, including politics, sports, international affairs, culture, weather, mythology, religion, science, and geography. The complete corpus contains 6,162 annotated sentences and 69,193 tokens, with a vocabulary of 14,349 unique tokens. Entity annotations include 4,609 PER, 2,904 ORG, 4,678 LOC, and 3,172 MISC tokens, alongside 53,830 non-entity (O) tokens.
All annotations follow the BIO tagging scheme and are provided in CoNLL format, making the dataset directly compatible with modern NLP frameworks such as Hugging Face Transformers, PyTorch, TensorFlow, and spaCy. In addition to the dataset, this repository includes the complete source code for data preprocessing. This dataset is intended for training, evaluation, benchmarking, and comparative research on Bangla NER systems, particularly transformer-based and hybrid sequence-labeling models.
It also serves as a valuable resource for downstream Bangla NLP applications, including information extraction, question answering, and knowledge graph construction.
Links
Where it is published
- DOI doi.org/10.17632/dy4nhv5tj2.2 ↗
DOI / persistent id · from researchdata tees ac uk
Catalogue records · 1
- OAI-PMH record data.mendeley.com/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Adata… ↗
metadata API · from researchdata tees ac uk
Topics
- From keywords
- Artificial intelligence · Computer Science & AI · Earth & Environmental Science · Engineering · Humanities · Life Sciences · Natural language processing · Social Science
- Inferred from text
- Text 75%
Provenance · 1 source records, 14 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Teesside University Research Data Repository | oai:data.mendeley.com/dy4nhv5tj2.2 | 7 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].anzsrc:field:460208 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Natural Language Processing'] |
| concepts[field].anzsrc:group:4602 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Artificial Intelligence'] |
| concepts[field].local:field:computer-science-ai | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:engineering | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:humanities | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:social-science | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[modality].local:modality:text | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/description |
| license | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/rights |
| publication_date | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| title | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/title |