Constarium
← Search

Table · dataset · 2022

Reddit-da

Listed in Hugging Face Datasets

Dataset Card for SQuAD-da Dataset

Description

Summary

This dataset consists of 1,908,887 Danish posts from Reddit. These are from this Reddit dump and have been filtered using this script, which uses FastText to detect the Danish posts. Supported Tasks and Leaderboards This dataset is suitable for language modelling.

Read the rest (1 more)

Languages This dataset is in Danish. Dataset Structure Data Instances Every entry in the dataset contains short Reddit… See the full description on the dataset page: huggingface.co/datasets/DDSC/reddit-da.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text · text generation
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsDDSC/reddit-da11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0