Data · dataset · 2024
To NER or not to NER? A case study of low-resource deontic modalities in EU legislation?
Listed in IISH Dataverse
Deontic modality (obligation, permission, prohibition) in legal documents can convey critical information, and identification of deontic modalities is often performed using Natural Language Processing (NLP) techniques as a `Deontic Modality Classification' (DMC) text classification task.
Description
As deontic modalities in legal text are not mutually exclusive, a key challenge with DMC is that it classifies the provided text into a single modality while in reality it might have multiple deontic modalities.
To address this, this study analyzes the feasibility of performing deontic modality identification as a Named Entity Recognition (NER) task over DMC task approaches in a low-resource data setting with EU legislation. Low-resource NLP approaches can offer solutions to tackle the problem of scarce data. In this paper, we use a rule-based approach with modal verbs and a Decision Tree classifier for DMC task.
Read the rest (1 more)
For NER, we utilize Conditional Random Fields (CRFs) in a low-resource setting and report on the reliability and precision for identification of deontic modality. Our experiments reveal that simpler models, like decision trees, out perform larger models in the low-resource setting of DMC obtaining macro-F1 score of 0.83. For the NER task, the CRF models show consistent performance for `obligation' labels with an F1-score of 0.51 but have wavering results for other classes with a max F1-score of 0.26 for `permission', and 0.08 for `prohibition'.
Links
Where it is published
- Dataverse dataset page dataverse.nl/dataset.xhtml?persistentId=doi%3A10.34894%2FD9AKUS ↗
landing page · from datasets iisg amsterdam
- DOI doi.org/10.34894/d9akus ↗
DOI / persistent id · from datasets iisg amsterdam
Catalogue records · 1
- Dataverse API dataverse.nl/api/datasets/:persistentId/?persistentId=doi%3A10.34894%2FD9AK… ↗
metadata API · from datasets iisg amsterdam
Topics
- Stated by source
- Computer and Information Science · Law
- From keywords
- Computer Science & AI · Humanities · Natural language processing · Social Science
- Inferred from text
- Text 75%
Provenance · 1 source records, 13 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| IISH Dataverse | doi:10.34894/D9AKUS | 8 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| concepts[field].anzsrc:field:460208 | mapping · datasets iisg amsterdam | vocabulary-mapper@1.0.0 | keywords['Natural Language Processing'] |
| concepts[field].dataverse_subject:computer-and-information-science | source · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | /subjects |
| concepts[field].dataverse_subject:law | source · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | /subjects |
| concepts[field].local:field:computer-science-ai | mapping · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | /subjects |
| concepts[field].local:field:humanities | mapping · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | /subjects |
| concepts[field].local:field:social-science | mapping · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | /subjects |
| concepts[modality].local:modality:text | enrichment · datasets iisg amsterdam | keyword-concept-rules@1.0.0 | title+description (75%) |
| created_date | source · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | |
| description | source · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | /description |
| publication_date | source · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | |
| title | source · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | /name |
| updated_date | source · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 | |
| version_label | source · datasets iisg amsterdam | connector:datasets_iisg_amsterdam@1.0.0 |