Constarium
← Search

Data · dataset · 2022

swahili

Listed in Hugging Face Datasets

The Swahili dataset developed specifically for language modeling task.

Description

The dataset contains 28,000 unique words with 6.84M, 970k, and 2M words for the train, valid and test partitions respectively which represent the ratio 80:10:10. The entire dataset is lowercased, has no punctuation marks and, the start and end of sentence markers have been incorporated to facilitate easy tokenization during language modeling.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
fill mask · text generation
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsuestc-swahili/swahili12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:fill-masksource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0