Constarium
← Search

Table · dataset · 2025

📄 FinePDFs

Listed in Hugging Face Datasets

Liberating 3T of the finest tokens from PDFs What is this?

Description

As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that.

It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: huggingface.co/datasets/HuggingFaceFW/finepdfs.

Links

Where it is published

Documentation and papers

Catalogue records · 1

Topics

Stated by source
tabular · text · text generation
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsHuggingFaceFW/finepdfs11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:tabularsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0