Data · dataset · 2022
classla/ssj500k
Listed in Hugging Face Datasets
The dataset contains 7432 training samples, 1164 validation samples and 893 test samples.
Description
Each sample represents a sentence and includes the following features: sentence ID ('sent_id'), list of tokens ('tokens'), list of lemmas ('lemmas'), list of Multext-East tags ('xpos_tags), list of UPOS tags ('upos_tags'), list of morphological features ('feats'), list of IOB tags ('iob_tags'), and list of universal dependency tags ('uds').
Three dataset configurations are available, where the corresponding features are encoded as class labels: 'ner', 'upos', and 'ud'.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/classla/ssj500k ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/classla/ssj500k ↗
metadata API · from Hugging Face
Topics
- Stated by source
- token classification
- From keywords
- Computer Science & AI
Provenance · 1 source records, 9 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | classla/ssj500k | 11 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[task].hf_task:token-classification | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| license | source · Hugging Face | connector:huggingface@1.0.0 | /tags[license:*] |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |