Constarium
← Search

Text · dataset · 2022

The Tatoeba Translation Challenge

Listed in Hugging Face Datasets

Dataset Card for The Tatoeba Translation Challenge Please note that this dataset is intended strictly for evaluation and benchmarking purposes.

Description

Training models on this dataset, or including it in automatically collected web-scale training corpora, may lead to benchmark contamination and invalidate evaluation results. Dataset

Summary

Read the rest (1 more)

The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed… See the full description on the dataset page: huggingface.co/datasets/Helsinki-NLP/tatoeba_mt.

Links

Get the data

Catalogue records · 1

Topics

Stated by source
text · text generation · translation
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsHelsinki-NLP/tatoeba_mt12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:translationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0