Constarium
← Search

Table · dataset · 2022

STSb Multi MT

Listed in Hugging Face Datasets

Dataset Card for STSb Multi MT Dataset

Description

Summary

STS Benchmark comprises a selection of the English datasets used in the STS tasks organized in the context of SemEval between 2012 and 2017. The selection of datasets include text from image captions, news headlines and user forums. (source) These are different multilingual translations and the English original of the STSbenchmark dataset.

Read the rest (1 more)

Translation has been done with deepl.com. It can be used to train sentence embeddings… See the full description on the dataset page: huggingface.co/datasets/PhilipMay/stsb_multi_mt.

Links

Where it is published

Documentation and papers

Catalogue records · 1

Topics

Stated by source
text · text classification
Inferred from text
Image 75% · Text 75%
Provenance · 1 source records, 12 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsPhilipMay/stsb_multi_mt11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:imageenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[task].hf_task:text-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0