Constarium
← Search

Table · dataset · 2022

tomekkorbak/pile-curse-full

Listed in Hugging Face Datasets

Description

Generation procedure The dataset was constructed using documents from the Pile scored using LDNOOBW wordlist (a score is number of curses per character). The procedure was the following: The first half of the data are 100k documents randomly sampled from the Pile and assigned scores The second half are the most cursing document from the Pile, obtained by scoring the whole Pile and choosing 100k documents with highest scores Then, the dataset was shuffled and a 9:1 train-test split… See the full description on the dataset page: huggingface.co/datasets/tomekkorbak/pile-curse-full.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text
Provenance · 1 source records, 8 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetstomekkorbak/pile-curse-full11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0