Constarium
← Search

Text · dataset · 2022

OffensEval 2020

Listed in Hugging Face Datasets

OffensEval 2020 features a multilingual dataset with five languages.

Description

The languages included in OffensEval 2020 are:

  • Arabic
  • Danish
  • English
  • Greek
  • Turkish The annotation follows the hierarchical tagset proposed in the Offensive Language Identification Dataset (OLID) and used in OffensEval 2019. In this taxonomy we break down offensive content into the following three sub-tasks taking the type and target of offensive content into account. The following sub-tasks were organized:
  • Sub-task A - Offensive language identification;
  • Sub-task B - Automatic categorization of offense types;
  • Sub-task C - Offense target identification. The English training data isn't included here (the text isn't available and needs rehydration of 9 million tweets; see https://zenodo.org/record/3950379#.XxZ-aFVKipp)

Links

Documentation and papers

Catalogue records · 1

Topics

Stated by source
text · text classification
Inferred from text
Text 75%
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsstrombergnlp/offenseval_202011 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[task].hf_task:text-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0