Constarium
← Search

Text · dataset · 2022

Catalan-German aligned corpora to train NMT systems.

Listed in Hugging Face Datasets

Dataset Card for Tilde-MODEL-Catalan Dataset

Description

Summary

This dataset contains the German version of the Tilde-MODEL corpus aligned with a Catalan translation. The catalan text has been obtained using Apertium's RBMT system from the Spanish version. It contains 3.4M segments.

Read the rest (1 more)

Supported Tasks and Leaderboards This dataset can be used to train NMT and SMT systems. It has been used as a training corpus for the Softcatalà machine translation engine. Languages… See the full description on the dataset page: huggingface.co/datasets/softcatala/Tilde-MODEL-Catalan.

Links

Catalogue records · 1

Topics

Stated by source
text · translation
Inferred from text
Text 75%
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetssoftcatala/Tilde-MODEL-Catalan12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[task].hf_task:translationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0