Table · dataset · 2022
Saibo-creator/bookcorpus_compact_1024
Listed in Hugging Face Datasets
Description
Dataset Card for "bookcorpus_compact_1024" Num samples: 616,051 The number of tokens for each sequence is not exactly 1024, but all slightly shorter than 1024. The sequences were built by merging sentences to the maximal length shorter than 1024 tokens. Therefore, padding is necessary for batch processing. import time from typing import List from datasets import load_dataset, Dataset from tqdm import tqdm from transformers import AutoTokenizer def batch_tokenize(texts: List[str]… See the full description on the dataset page: huggingface.co/datasets/Saibo-creator/bookcorpus_compact_1024.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/Saibo-creator/bookcorpus_compact_1024 ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/Saibo-creator/bookcorpus_compact_1024 ↗
metadata API · from Hugging Face
Topics
- Stated by source
- text
- From keywords
- Computer Science & AI
Provenance · 1 source records, 8 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | Saibo-creator/bookcorpus_compact_1024 | 12 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |