Constarium
← Search

Data · dataset · 2023

Intrinsic Evaluation Dataset for Automatic Sentence Alignment

Listed in ScienceDB

Description

This dataset is for the intrinsic evaluation of automatic sentence alignment. It consists of manually aligned sentence pairs from the publicly available German-French parallel corpus Text+Berg (github.com/rsennrich/Bleualign/tree/master/eval) and two newly built corpora based on the English, German, French, Japanese and Italian translations of the Chinese novel The Three-Body Problem and political documents selected from The Governance of China.

All the sentence pairs in the dataset have been checked by human annotators, which can be used as the gold standard to evaluate the output of automatic sentence aligners.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Linguistics 71% · Text 75%
Provenance · 1 source records, 12 field assertions
SourceKeyLast seenRaw
ScienceDB10.57760/sciencedb.j00133.003218 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · scidb cnconnector:scidb_cn@1.0.0
concepts[field].anzsrc:group:4704enrichment · scidb cntaxonomy-embedding@1.0.0title+keywords+description (71%)
concepts[field].local:field:earth-environmentalmapping · scidb cnconnector:scidb_cn@1.0.0
concepts[field].local:field:engineeringmapping · scidb cnconnector:scidb_cn@1.0.0
concepts[field].local:field:humanitiesmapping · scidb cnconnector:scidb_cn@1.0.0
concepts[field].local:field:life-sciencesmapping · scidb cnconnector:scidb_cn@1.0.0
concepts[field].local:field:social-sciencemapping · scidb cnconnector:scidb_cn@1.0.0
concepts[modality].local:modality:textenrichment · scidb cnkeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · scidb cnconnector:scidb_cn@1.0.0/metadata/dc/description
license_textsource · scidb cnconnector:scidb_cn@1.0.0
publication_datesource · scidb cnconnector:scidb_cn@1.0.0
titlesource · scidb cnconnector:scidb_cn@1.0.0/metadata/dc/title