Table · dataset · 2025
Dargwa Monolingual Corpus
Listed in Hugging Face Datasets
Description
📖 Dargwa Monolingual Corpus The Dargwa Monolingual Corpus consists of texts in the Dargwa language, collected from various sources. This dataset is useful for tasks such as language modeling, tokenization, and linguistic research. Below is a concise description of the dataset columns: Columns Description id – Unique identifier for each entry. source – URL or source identifier from where the text was extracted. language – Language of the text (always "dar" for this… See the full description on the dataset page: huggingface.co/datasets/Murtazali/dargwa-monocorpus.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/Murtazali/dargwa-monocorpus ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/Murtazali/dargwa-monocorpus ↗
metadata API · from Hugging Face
Topics
- Stated by source
- text
- From keywords
- Computer Science & AI · Text
Provenance · 1 source records, 9 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | Murtazali/dargwa-monocorpus | 12 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].local:modality:text | mapping · Hugging Face | vocabulary-mapper@1.0.0 | keywords['Text Corpus'] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |