Constarium
← Search

Table · dataset · 2022

Berlin State Library OCR

Listed in Hugging Face Datasets

Dataset Card for Berlin State Library OCR data Dataset

Description

Summary

The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945. At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages. For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012).

Read the rest (1 more)

Supported Tasks and Leaderboards language-modeling: this dataset has the potential to be used… See the full description on the dataset page: huggingface.co/datasets/SBB/sbb-dc-ocr.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
fill mask · text · text generation
Inferred from text
Text 75%
Provenance · 1 source records, 13 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsSBB/sbb-dc-ocr10 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[method].local:method:ocrmapping · Hugging Facevocabulary-mapper@1.0.0keywords['ocr']
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[task].hf_task:fill-masksource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0