Constarium
← Search

Text · dataset · 2023

CC OpenBooks

Listed in Hugging Face Datasets

Dataset Card for CC OpenBooks Dataset Description CC OpenBooks is a curated collection of high quality non-fiction books.

Description

All texts are from CC-By-4.0 sources, with no license ambiguity. The documents are normalized to markdown, and care is taken to ensure most formatting (e.g. inline LaTeX) remains intact.

Files are manually inspected and cleaned of all defects wherever possible. Source Data The following Openstax collections were used in creating this… See the full description on the dataset page: huggingface.co/datasets/Daniel-P-Gonzalez/CCOpenBooks.

Links

Documentation and papers

Catalogue records · 1

Topics

Stated by source
text · text generation
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsDaniel-P-Gonzalez/CCOpenBooks12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0