Table · dataset · 2022
Reddit-da
Listed in Hugging Face Datasets
Dataset Card for SQuAD-da Dataset
Description
Summary
This dataset consists of 1,908,887 Danish posts from Reddit. These are from this Reddit dump and have been filtered using this script, which uses FastText to detect the Danish posts. Supported Tasks and Leaderboards This dataset is suitable for language modelling.
Read the rest (1 more)
Languages This dataset is in Danish. Dataset Structure Data Instances Every entry in the dataset contains short Reddit… See the full description on the dataset page: huggingface.co/datasets/DDSC/reddit-da.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/DDSC/reddit-da ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/DDSC/reddit-da ↗
metadata API · from Hugging Face
Topics
- Stated by source
- text · text generation
- From keywords
- Computer Science & AI
Provenance · 1 source records, 10 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | DDSC/reddit-da | 11 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[task].hf_task:text-generation | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| license | source · Hugging Face | connector:huggingface@1.0.0 | /tags[license:*] |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |