Constarium
← Search

Data · dataset · 2022

classla/ssj500k

Listed in Hugging Face Datasets

The dataset contains 7432 training samples, 1164 validation samples and 893 test samples.

Description

Each sample represents a sentence and includes the following features: sentence ID ('sent_id'), list of tokens ('tokens'), list of lemmas ('lemmas'), list of Multext-East tags ('xpos_tags), list of UPOS tags ('upos_tags'), list of morphological features ('feats'), list of IOB tags ('iob_tags'), and list of universal dependency tags ('uds').

Three dataset configurations are available, where the corresponding features are encoded as class labels: 'ner', 'upos', and 'ud'.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
token classification
Provenance · 1 source records, 9 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsclassla/ssj500k11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:token-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0