Constarium
← Search

Text · dataset · 2022

BLURB

Listed in Hugging Face Datasets

The BioCreative II Gene Mention task.

Description

The training corpus for the current task consists mainly of the training and testing corpora (text collections) from the BCI task, and the testing corpus for the current task consists of an additional 5,000 sentences that were held 'in reserve' from the previous task. In the current corpus, tokenization is not provided; instead participants are asked to identify a gene mention in a sentence by giving its start and end characters.

As before, the training set consists of a set of sentences, and for each sentence a set of gene mentions (GENE annotations). - Homepage: biocreative.bioinformatics.udel.edu/tasks/biocreative-ii/task-1a-gene-mention-tagging/ - Repository: github.com/cambridgeltl/MTL-Bioinformatics-2016/raw/master/data/ - Paper: Overview of BioCreative II gene mention recognition link.springer.com/article/10.1186/gb-2008-9-s2-s2

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text
Inferred from text
Text 75%
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsbigbio/blurb10 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0