Constarium
← Search

Table · dataset · 2025

Dargwa Monolingual Corpus

Listed in Hugging Face Datasets

Description

📖 Dargwa Monolingual Corpus The Dargwa Monolingual Corpus consists of texts in the Dargwa language, collected from various sources. This dataset is useful for tasks such as language modeling, tokenization, and linguistic research. Below is a concise description of the dataset columns: Columns Description id – Unique identifier for each entry. source – URL or source identifier from where the text was extracted. language – Language of the text (always "dar" for this… See the full description on the dataset page: huggingface.co/datasets/Murtazali/dargwa-monocorpus.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text
From keywords
Computer Science & AI · Text
Provenance · 1 source records, 9 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsMurtazali/dargwa-monocorpus12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textmapping · Hugging Facevocabulary-mapper@1.0.0keywords['Text Corpus']
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0