Data · dataset · 2026
German News Portals: The Linked Domains In Their Articles
Listed in TUDOdata
The dataset was created to investigate the linking behavior of 12 of the largest German news portals, with the goal of understanding how they connect to other online resources.
Description
Specifically, this study aims to examine whether the number of links correlates with other variables such as article title, keywords, section, and length. By analyzing these relationships, we hope to gain insights into the factors that influence the linking behavior of news portals.
The data was automatically scraped using a web crawler implemented in Python. For copyright reasons, the article texts are not included.
Read the rest (7 more)
Methods
To collect the data, we employed a web crawling and scraping approach, where the homepages of the 12 news portals were crawled daily at 12 noon during the survey period (from 2024-08-27 until 2024-10-31). We used the Scrapy and Selenium libraries in Python to implement the crawler and scraper, which allowed us to handle cookie banners, interactive elements, and dynamic content when necessary. The crawler extracted all news articles in text form on the homepage, along with the links contained within the article text.
The code is available on Github at github.com/EstherKuerbis/NewsArticlesScraper.git, allowing other researchers to review, use, and modify the code for their own purposes. To exclude irrelevant links, we filtered out links that appeared alongside the article text, such as advertisements. Where possible, we extracted the articles’ departments from the scraped keywords or URLs.
To categorize the domains, we used a Python script to automatically subdivide them into internal and external domains based on the domain name in the URL. This approach enabled us to collect a comprehensive dataset of news articles and links, which can be used to analyze the linking behavior of news portals. During the survey period (from 2024-08-27 until 2024-10-31), the homepages of the news portals were crawled daily at 12 noon.
All news articles in text form on the homepage were also crawled. The article texts and the links contained therein were scraped. Links that appeared alongside the article text (e.g. advertisements) were excluded.
The crawler and scraper were implemented in Python using the Scrapy and Selenium libraries. This made it possible to handle cookie banners or other interactive and dynamic elements when necessary. Where possible, the articles' departments were extracted from the scraped keywords or the URLs.
The domains were automatically subdivided into internal and external using a Python script based on the domain name in the URL.
Links
Where it is published
- Dataverse dataset page data.tu-dortmund.de/dataset.xhtml?persistentId=doi%3A10.71955%2FDUEDATA-2026-ML5CA… ↗
landing page · from data tu dortmund de
- DOI doi.org/10.71955/duedata-2026-ml5catjg ↗
DOI / persistent id · from data tu dortmund de
Catalogue records · 1
- Dataverse API data.tu-dortmund.de/api/datasets/:persistentId/?persistentId=doi%3A10.71955%2FDUED… ↗
metadata API · from data tu dortmund de
Topics
- Stated by source
- Computer and Information Science
- From keywords
- Chemistry · Engineering · Humanities · Mathematics & Statistics · Psychology & Behavioral Science · Social Science
- Inferred from text
- Data management and data science 68% · Text 75%
Provenance · 1 source records, 15 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| TUDOdata | doi:10.71955/DUEDATA-2026-ML5CATJG | 8 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| concepts[field].anzsrc:group:4605 | enrichment · data tu dortmund de | taxonomy-embedding@1.0.0 | title+keywords+description (68%) |
| concepts[field].dataverse_subject:computer-and-information-science | source · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /subjects |
| concepts[field].local:field:chemistry | mapping · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /subjects |
| concepts[field].local:field:engineering | mapping · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /subjects |
| concepts[field].local:field:humanities | mapping · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /subjects |
| concepts[field].local:field:mathematics-statistics | mapping · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /subjects |
| concepts[field].local:field:psychology-behavioral | mapping · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /subjects |
| concepts[field].local:field:social-science | mapping · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /subjects |
| concepts[modality].local:modality:text | enrichment · data tu dortmund de | keyword-concept-rules@1.0.0 | title+description (75%) |
| created_date | source · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | |
| description | source · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /description |
| publication_date | source · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | |
| title | source · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | /name |
| updated_date | source · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 | |
| version_label | source · data tu dortmund de | connector:data_tu_dortmund_de@1.0.0 |