Data · dataset · 2018
Annotated Corpus For Occitan
Listed in DataCite
This corpus contains a collection of texts in Occitan which were manually annotated with parts-of-speech, lemmas.
Description
The corpus was produced in the context of the RESTAURE project, funded by the French ANR. The current version of the corpus contains 28 documents and 12,425 tokens.
The annotation process is detailed in the following article: hal.archives-ouvertes.fr/hal-01704806 The annotated versions are provided in a TSV CoNLL-U format.
Links
Where it is published
- Repository landing page zenodo.org/record/1182949 ↗
landing page · from DataCite
- DOI doi.org/10.5281/zenodo.1182949 ↗
DOI / persistent id · from DataCite
Documentation and papers
- Creative Commons Attribution Share-Alike 4.0 creativecommons.org/licenses/by-sa/4.0 ↗
license · from DataCite
Catalogue records · 2
- DataCite API api.datacite.org/dois/10.5281/zenodo.1182949 ↗
metadata API · from DataCite
- DataCite Commons commons.datacite.org/doi.org/10.5281/zenodo.1182949 ↗
catalogue entry · from DataCite
Topics
- Stated by source
- Languages and literature
Related
- Newer version ofAnnotated Corpus For Occitan
- Inverse of has versionAnnotated Corpus For Occitan
Provenance · 1 source records, 6 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| DataCite | 10.5281/zenodo.1182949 | 12 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · DataCite | connector:datacite@1.0.0 | /data/attributes/rightsList |
| concepts[field].fos:languages-and-literature | source · DataCite | connector:datacite@1.0.0 | |
| description | source · DataCite | connector:datacite@1.0.0 | /data/attributes/descriptions |
| license | source · DataCite | connector:datacite@1.0.0 | /data/attributes/rightsList |
| publication_date | source · DataCite | connector:datacite@1.0.0 | /data/attributes/dates |
| title | source · DataCite | connector:datacite@1.0.0 | /data/attributes/titles/0/title |