Constarium
← Search

Data · dataset · 2025

ACE 2005 Multilingual Training Corpus

Listed in UC Berkeley Library Dataverse

ACE 2005 Multilingual Training Corpus was developed by the Linguistic Data Consortium (LDC) and contains approximately 1,800 files of mixed genre text in English, Arabic, and Chinese annotated for entities, relations, and events.

Description

This represents the complete set of training data in those languages for the 2005 Automatic Content Extraction (ACE) technology evaluation. The genres include newswire, broadcast news, broadcast conversation, weblog, discussion forums, and conversational telephone speech.

The data was annotated by LDC with support from the ACE Program and additional assistance from LDC. The objective of the ACE program was to develop automatic content extraction technology to support automatic processing of human language in text form. In November 2005, sites were evaluated on system performance in five primary areas: the recognition of entities, values, temporal expressions, relations, and events.

Read the rest (3 more)

Entity, relation, and event mention detection were also offered as diagnostic tasks. All tasks with the exception of event tasks were performed for three languages, English, Chinese, and Arabic. Events tasks were evaluated in English and Chinese only.

This release comprises the official training data for these evaluation tasks. For more information about linguistic resources for the ACE Program, including annotation guidelines, task definitions and other documentation, see LDC's ACE website. Suggested citation: Walker, Christopher, et al. ACE 2005 Multilingual Training Corpus LDC2006T06.

Web Download. Philadelphia: Linguistic Data Consortium, 2006.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
Social Sciences
From keywords
Social Science
Inferred from text
Audio 65% · Text 75%
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
UC Berkeley Library Dataversedoi:10.60503/D3/CDCBX15 d agoJSON v1
FieldAssertionExtractorEvidence
concepts[field].dataverse_subject:social-sciencessource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0/subjects
concepts[field].local:field:social-sciencemapping · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0/subjects
concepts[modality].local:modality:audioenrichment · datasets lib berkeley edukeyword-concept-rules@1.0.0title+description (65%)
concepts[modality].local:modality:textenrichment · datasets lib berkeley edukeyword-concept-rules@1.0.0title+description (75%)
created_datesource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0
descriptionsource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0/description
publication_datesource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0
titlesource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0/name
updated_datesource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0
version_labelsource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0