Text · dataset · 2022
PlanTL-GOB-ES/pharmaconer
Listed in Hugging Face Datasets
PharmaCoNER: Pharmacological Substances, Compounds and Proteins Named Entity Recognition track This dataset is designed for the PharmaCoNER task, sponsored by Plan de Impulso de las Tecnologías del Lenguaje (Plan TL).
Description
It is a manually classified collection of clinical case studies derived from the Spanish Clinical Case Corpus (SPACCC), an open access electronic library that gathers Spanish medical publications from SciELO (Scientific Electronic Library Online).
The annotation of the entire set of entity mentions was carried out by medicinal chemistry experts and it includes the following 4 entity types: NORMALIZABLES, NO_NORMALIZABLES, PROTEINAS and UNCLEAR. The PharmaCoNER corpus contains a total of 396,988 words and 1,000 clinical cases that have been randomly sampled into 3 subsets. The training set contains 500 clinical cases, while the development and test sets contain 250 clinical cases each.
Read the rest (1 more)
In terms of training examples, this translates to a total of 8074, 3764 and 3931 annotated sentences in each set. The original dataset was distributed in Brat format (brat.nlplab.org/standoff.html). For further information, please visit temu.bsc.es/pharmaconer/ or send an email to encargo-pln-life@bsc.es
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/PlanTL-GOB-ES/pharmaconer ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/PlanTL-GOB-ES/pharmaconer ↗
metadata API · from Hugging Face
Topics
- Stated by source
- text · token classification
- From keywords
- Computer Science & AI
Provenance · 1 source records, 10 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | PlanTL-GOB-ES/pharmaconer | 11 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[task].hf_task:token-classification | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| license | source · Hugging Face | connector:huggingface@1.0.0 | /tags[license:*] |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |