Data · dataset · 2026
News on the Web
Listed in UC Berkeley Library Dataverse
The NOW corpus (News on the Web) has been created by Mark Davies, and it contains 25.8 billion words of data from web-based newspapers and magazines from 2010 to the present time.
Description
More importantly, the corpus grows by about 240-250 million words of data each month (from about 400,000 new articles), or about three billion words each year. While other resources like Google Trends show you what people are searching for, the NOW Corpus is the only structured corpus that shows you what is actually happening in the language -- virtually right up to the present time.
For example, see the frequency of words since 2010, as well as new words and phrases from the last few years. In this sense, NOW is the most robust monitor corpus of English. Finally, the corpus is related to other corpora from English-Corpora.org, which are the most widely used corpora of English and which offer unparalleled insight into variation in English.
Read the rest (1 more)
The dataset in UC Berkeley Library's Dataverse is available to download or users can query the data by visiting: NOW Corpus News on the Web.
Links
Where it is published
- Dataverse dataset page datasets.lib.berkeley.edu/dataset.xhtml?persistentId=doi%3A10.60503%2FD3%2FPLBI9D ↗
landing page · from datasets lib berkeley edu
- DOI doi.org/10.60503/d3/plbi9d ↗
DOI / persistent id · from datasets lib berkeley edu
Catalogue records · 1
- Dataverse API datasets.lib.berkeley.edu/api/datasets/:persistentId/?persistentId=doi%3A10.60503%2FD3%2… ↗
metadata API · from datasets lib berkeley edu
Topics
- Stated by source
- Arts and Humanities · Social Sciences
- From keywords
- Humanities · Social Science
- Inferred from text
- Linguistics 70%
Provenance · 1 source records, 11 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| UC Berkeley Library Dataverse | doi:10.60503/D3/PLBI9D | 4 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| concepts[field].anzsrc:group:4704 | enrichment · datasets lib berkeley edu | taxonomy-embedding@1.1.0 | title+keywords+description (70%) |
| concepts[field].dataverse_subject:arts-and-humanities | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /subjects |
| concepts[field].dataverse_subject:social-sciences | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /subjects |
| concepts[field].local:field:humanities | mapping · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /subjects |
| concepts[field].local:field:social-science | mapping · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /subjects |
| created_date | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | |
| description | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /description |
| publication_date | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | |
| title | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /name |
| updated_date | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | |
| version_label | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 |