Data · dataset · 2017
Learning multilingual named entity recognition from Wikipedia
Listed in DataCite
Description
This is the data associated with Joel Nothman, Nicky Ringland, Will Radford, Tara Murphy and James R. Curran (2013), "Learning multilingual named entity recognition from Wikipedia", Artificial Intelligence 194 (DOI: 10.1016/j.artint.2012.03.006). A preprint is included here as wikiner-preprint.pdf This data was originally available at schwa.org/resources (which linked to schwa.org/projects/resources/wiki/Wikiner).
The .bz2 files are NER training corpora produced as reported in the Artificial Intelligence paper. wp2 and wp3 are differentiated by wp3 using a higher level of link inference. They use a pipe-delimited format that can be converted to CoNLL 2003 format with system2conll.pl. nothman08types.tsv is a manual classification of articles first used in Joel Nothman, James R. Curran and Tara Murphy (2008), "Transforming Wikipedia into Named Entity Training Data", In Proceedings of the Australasian Language Technology Association Workshop 2008 . aclanthology.coli.uni-saarland.de/pdf/U/U08/U08-1016.pdf popular.tsv and random.tsv are manual article classifications developed for the Artifiical Intelligence paper based on different strategies for sampling articles from Wikipedia in order to account for Wikipedia's biased distribution (see that paper). scheme.tsv maps these fine-grained labels to coarser annotations including CoNLL 2003-style. wikigold.conll.txt is a manual NER annotation of some Wikipedia text as presented in Dominic Balasuriya and Nicky Ringland and Joel Nothman and Tara Murphy and James R. Curran (2009), in Proceedings of the 2009 Workshop on The People's Web Meets NLP: Collaboratively Constructed Semantic Resources (aclweb.org/anthology/W/W09/W09-3302).
Read the rest (1 more)
See also corpora produced similarly in an enhanced version of this work work (Pan et al., "Cross-lingual Name Tagging and Linking for 282 Languages", ACL 2017) at nlp.cs.rpi.edu/wikiann/.
Links
Where it is published
- Repository landing page figshare.com/articles/dataset/Learning_multilingual_named_entity_recognitio… ↗
landing page · from DataCite
- DOI doi.org/10.6084/m9.figshare.5462500.v1 ↗
DOI / persistent id · from DataCite
Documentation and papers
Catalogue records · 2
- DataCite API api.datacite.org/dois/10.6084/m9.figshare.5462500.v1 ↗
metadata API · from DataCite
- DataCite Commons commons.datacite.org/doi.org/10.6084/m9.figshare.5462500.v1 ↗
catalogue entry · from DataCite
Topics
- Stated by source
- Computer and information sciences · Languages and literature · Psychology
- Inferred from text
- Text 75%
Related
Provenance · 1 source records, 12 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| DataCite | 10.6084/m9.figshare.5462500.v1 | 12 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · DataCite | connector:datacite@1.0.0 | /data/attributes/rightsList |
| byte_size | source · DataCite | connector:datacite@1.0.0 | |
| concepts[field].fos:computer-and-information-sciences | source · DataCite | connector:datacite@1.0.0 | |
| concepts[field].fos:languages-and-literature | source · DataCite | connector:datacite@1.0.0 | |
| concepts[field].fos:psychology | source · DataCite | connector:datacite@1.0.0 | |
| concepts[modality].local:modality:text | enrichment · DataCite | keyword-concept-rules@1.0.0 | title+description (75%) |
| created_date | source · DataCite | connector:datacite@1.0.0 | |
| description | source · DataCite | connector:datacite@1.0.0 | /data/attributes/descriptions |
| license | source · DataCite | connector:datacite@1.0.0 | /data/attributes/rightsList |
| publication_date | source · DataCite | connector:datacite@1.0.0 | /data/attributes/dates |
| title | source · DataCite | connector:datacite@1.0.0 | /data/attributes/titles/0/title |
| updated_date | source · DataCite | connector:datacite@1.0.0 |