Constarium
← Search

Data · dataset · 2022

n2c2 2006 Smoking Status

Listed in Hugging Face Datasets

The data for the n2c2 2006 smoking challenge consisted of discharge summaries from Partners HealthCare, which were then de-identified, tokenized, broken into sentences, converted into XML format, and separated into training and test sets.

Description

Two pulmonologists annotated each record with the smoking status of patients based strictly on the explicitly stated smoking-related facts in the records. These annotations constitute the textual judgments of the annotators.

The annotators were asked to classify patient records into five possible smoking status categories: a past smoker, a current smoker, a smoker, a non-smoker and an unknown. A total of 502 de-identified medical discharge records were used for the smoking challenge.

Links

Where it is published

Catalogue records · 1

Topics

Provenance · 1 source records, 8 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsbigbio/n2c2_2006_smokers12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0