Constarium
← Search

Data · dataset · 2016

Reddit Cross-Topic Authorship Verification Corpus

Listed in DataCite

The "Reddit Cross-Topic Authorship Verification Corpus" consists of comments written between 2010 to 2016 from 1,000 reddit users.

Description

Each problem includes 1 unknown and 4 known documents (~ 7 KByte per document), where each document represents an aggregation of reviews coined from the same so-called subreddit. More precisely, all documents within a problem are disjunct regarding the subreddits in order to enable a cross-topic corpus.

All subreddits cover exactly 1,388 different topics such as books, news, gaming, music, movies, etc. The corpus follows excatly the same format as the well-known PAN Authorship Identification corpora (pan.webis.de/).

Links

Where it is published

Documentation and papers

Catalogue records · 2

Topics

Stated by source
Languages and literature
Provenance · 1 source records, 7 field assertions
SourceKeyLast seenRaw
DataCite10.17632/hppkn5kbg8.110 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · DataCiteconnector:datacite@1.0.0/data/attributes/rightsList
concepts[field].fos:languages-and-literaturesource · DataCiteconnector:datacite@1.0.0
concepts[field].local:field:computer-science-aimapping · DataCitevocabulary-mapper@1.0.0keywords['Machine Learning']
descriptionsource · DataCiteconnector:datacite@1.0.0/data/attributes/descriptions
licensesource · DataCiteconnector:datacite@1.0.0/data/attributes/rightsList
publication_datesource · DataCiteconnector:datacite@1.0.0/data/attributes/dates
titlesource · DataCiteconnector:datacite@1.0.0/data/attributes/titles/0/title