Constarium
← Search

Table · dataset · 2022

OpusTedtalks

Listed in Hugging Face Datasets

Dataset Card for OpusTedtalks Dataset

Description

Summary

This is a Croatian-English parallel corpus of transcribed and translated TED talks, originally extracted from wit3.fbk.eu. The corpus is compiled by Željko Agić and is taken from lt.ffzg.hr/zagic provided under the CC-BY-NC-SA license. This corpus is sentence aligned for both language pairs.

Read the rest (1 more)

The documents were collected and aligned using the Hunalign algorithm. Supported Tasks and Leaderboards… See the full description on the dataset page: huggingface.co/datasets/Helsinki-NLP/opus_tedtalks.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text · translation
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsHelsinki-NLP/opus_tedtalks12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:translationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0