Constarium
← Search

Table · dataset · 2022

bigscience-data/roots_ca_catalan_government_crawling

Listed in Hugging Face Datasets

ROOTS Subset: roots_ca_catalan_government_crawling Catalan Government Crawling Dataset uid: catalan_government_crawling Description The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web.

Description

It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39.117.909 tokens, 1.565.433 sentences and 71.043 documents.

Documents are separated… See the full description on the dataset page: huggingface.co/datasets/bigscience-data/roots_ca_catalan_government_crawling.

Links

Get the data

Catalogue records · 1

Topics

Stated by source
text
Provenance · 1 source records, 9 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsbigscience-data/roots_ca_catalan_government_crawling12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0