Constarium
← Search

Data · dataset · 2018

Annotated Corpus For Occitan

Listed in DataCite

This corpus contains a collection of texts in Occitan which were manually annotated with parts-of-speech, lemmas.

Description

The corpus was produced in the context of the RESTAURE project, funded by the French ANR. The current version of the corpus contains 28 documents and 12,425 tokens.

The annotation process is detailed in the following article: hal.archives-ouvertes.fr/hal-01704806 The annotated versions are provided in a TSV CoNLL-U format.

Links

Where it is published

Catalogue records · 2

Topics

Stated by source
Languages and literature

Related

Provenance · 1 source records, 6 field assertions
SourceKeyLast seenRaw
DataCite10.5281/zenodo.118294912 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · DataCiteconnector:datacite@1.0.0/data/attributes/rightsList
concepts[field].fos:languages-and-literaturesource · DataCiteconnector:datacite@1.0.0
descriptionsource · DataCiteconnector:datacite@1.0.0/data/attributes/descriptions
licensesource · DataCiteconnector:datacite@1.0.0/data/attributes/rightsList
publication_datesource · DataCiteconnector:datacite@1.0.0/data/attributes/dates
titlesource · DataCiteconnector:datacite@1.0.0/data/attributes/titles/0/title