Constarium
← Search

Table · dataset · 2023

UzBooks

Listed in Hugging Face Datasets

Dataset Card for BookCorpus Dataset

Description

Summary

In an effort to democratize research on low-resource languages, we release UzBooks dataset, a cleaned book corpus consisting of nearly 40000 books in Uzbek Language divided into two branches: "original" and "lat," representing the OCRed (Latin and Cyrillic) and fully Latin versions of the texts, respectively. Please refer to our blogpost and paper (Coming soon!) for further details. To load and use dataset, run this script:… See the full description on the dataset page: huggingface.co/datasets/murodbek/uz-books.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
fill mask · text · text generation
Inferred from text
Library and information studies 69%
Provenance · 1 source records, 12 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsmurodbek/uz-books8 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].anzsrc:group:4610enrichment · Hugging Facetaxonomy-embedding@1.1.0title+keywords+description (69%)
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:fill-masksource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0