Constarium
← Search

Table · dataset · 2022

wiki-paragraphs

Listed in Hugging Face Datasets

Dataset Card for wiki-paragraphs Dataset

Description

Summary

The wiki-paragraphs dataset is constructed by automatically sampling two paragraphs from a Wikipedia article. If they are from the same section, they will be considered a "semantic match", otherwise as "dissimilar". Dissimilar paragraphs can in theory also be sampled from other documents, but have not shown any improvement in the particular evaluation of the linked work.The alignment is in no way meant as an accurate… See the full description on the dataset page: huggingface.co/datasets/dennlinger/wiki-paragraphs.

Links

Where it is published

Documentation and papers

Catalogue records · 1

Topics

Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsdennlinger/wiki-paragraphs10 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:sentence-similaritysource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:text-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0