Constarium
← Search

Text · dataset · 2022

dataset/wikipedia_bn

Listed in Hugging Face Datasets

Description

Bengali Wikipedia from the dump of 03/20/2021. The data was processed using the huggingface datasets wikipedia script early april 2021. The dataset was built from the Wikipedia dump (dumps.wikimedia.org/).

Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text
Provenance · 1 source records, 8 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsdataset/wikipedia_bn11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0