Data · dataset · 2025
Explainable Detecting Noncompliance in Privacy Agreement based on Large Language Model Fine-tune-dataset
Listed in ScienceDB
The dataset includes three sub datasets of the article, all of which are datasets used for model training in the article "Explanatory Detection of Privacy Protocol Violations Based on Large Language Model Fine tuning".
Description
This includes three parts: a dataset for pre classifying privacy protocol texts, a named entity recognition corpus for identifying key content of privacy protocols, and a corpus dataset for detecting violations using a large language model. The dataset for pre classification of privacy protocol text contains a total of 15627 data items, including classification labels and corresponding text data that have been classified according to privacy protocols. The original annotated corpus for named entity recognition of key content in privacy protocols includes data annotated with BOEM for named entity recognition of privacy protocols for 40 apps. The corpus dataset for violation detection using the big language model consists of three parts.
The first part is the common knowledge of regulations, which includes a total of 147 basic contents from the Personal Information Security Specification; The second part is compliance labeling, which includes a training dataset of 1488 data sets obtained by balancing the core content of the privacy agreement with violation labeling corpora; The third part is the semantic interpretation text of the regulations, which includes a dataset of 139 articles obtained by manually interpreting some of the content in the second part of the data.
Links
Where it is published
- DOI doi.org/10.57760/sciencedb.27244 ↗
DOI / persistent id · from scidb cn
Catalogue records · 1
- OAI-PMH record scidb.cn/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=10.57760%2… ↗
metadata API · from scidb cn
Topics
- From keywords
- Earth & Environmental Science · Engineering · Humanities · Life Sciences · Social Science
- Inferred from text
- Text 75%
Provenance · 1 source records, 10 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| ScienceDB | 10.57760/sciencedb.27244 | 9 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| concepts[field].local:field:earth-environmental | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:engineering | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:humanities | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:social-science | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[modality].local:modality:text | enrichment · scidb cn | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · scidb cn | connector:scidb_cn@1.0.0 | /metadata/dc/description |
| license_text | source · scidb cn | connector:scidb_cn@1.0.0 | |
| publication_date | source · scidb cn | connector:scidb_cn@1.0.0 | |
| title | source · scidb cn | connector:scidb_cn@1.0.0 | /metadata/dc/title |