Table · dataset · 2022
tomekkorbak/pile-curse-full
Listed in Hugging Face Datasets
Description
Generation procedure The dataset was constructed using documents from the Pile scored using LDNOOBW wordlist (a score is number of curses per character). The procedure was the following: The first half of the data are 100k documents randomly sampled from the Pile and assigned scores The second half are the most cursing document from the Pile, obtained by scoring the whole Pile and choosing 100k documents with highest scores Then, the dataset was shuffled and a 9:1 train-test split… See the full description on the dataset page: huggingface.co/datasets/tomekkorbak/pile-curse-full.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/tomekkorbak/pile-curse-full ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/tomekkorbak/pile-curse-full ↗
metadata API · from Hugging Face
Topics
- Stated by source
- text
- From keywords
- Computer Science & AI
Provenance · 1 source records, 8 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | tomekkorbak/pile-curse-full | 11 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |