Constarium
← Search

Data · dataset · 2022

cdminix/iwslt2011

Listed in Hugging Face Datasets

Both manual transcripts and ASR outputs from the IWSLT2011 speech translation evalutation campaign are often used for the related punctuation annotation task.

Description

This dataset takes care of preprocessing said transcripts and automatically inserts punctuation marks given in the manual transcripts in the ASR outputs using Levenshtein aligment.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Artificial intelligence 75% · Audio 65%
Provenance · 1 source records, 9 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetscdminix/iwslt201112 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].anzsrc:group:4602enrichment · Hugging Facetaxonomy-embedding@1.1.0title+keywords+description (75%)
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:audioenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (65%)
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0