Hugging Face Datasets2022 · Table · Parquet
Berlin State Library OCR…At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.… …For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012).…
Hugging Face Datasets2022 · Image
Europeana NewspapersDataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiativ
Hugging Face Datasets2022 · Image
chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
Hugging Face Datasets2022 · dataset
ai-forever/Peter…The dataset annotation contain end-to-end markup for training detection and OCR models, as well as an end-to-end model for reading text from pages.…
Hugging Face Datasets2022 · dataset
ai-forever/school_notebooks_RU…The dataset annotation contain end-to-end markup for training detection and OCR models, as well as an end-to-end model for reading text from pages.…
Hugging Face Datasets2022 · Image
ai-forever/school_notebooks_EN…The dataset annotation contain end-to-end markup for training detection and OCR models, as well as an end-to-end model for reading text from pages.…
INESC TEC Research Data Repository2022 · Table · CSV
Manual Transcriptions of Typewritten Digital Representations of Portuguese Cultural Heritage Documents from the 20th CenturyThe dataset includes manual transcriptions of typewritten digital representations of Portuguese cultural heritage documents from the 20th century, extracted from the Arquivo Nacional Torre do Tombo (ANTT) (https://digitarq.arquivos.pt/). It has records from two funds: the General Administration of National Treasury (DGFP) and the National Secretariat of Information (SNI). The manual transcriptions
INESC TEC Research Data Repository2022 · Table · CSV
Typewritten Digital Representations of Portuguese Cultural Heritage Documents from the 20th century…Identifying digital representation typologies relevant to the OCR task was carried out by observing existing textual descriptions of the records and document layouts of the digital…
Hugging Face Datasets2022 · Image
SROIE_2019_text_recognition…This dataset we prepared using the Scanned receipts OCR and information extraction(SROIE) dataset. The SROIE dataset contains 973 scanned receipts in English language.…
Hugging Face Datasets2022 · Image
hf-internal-testing/fixtures_ocr…This dataset includes 2 images: one of the IAM Handwriting Database and one of the SRIOE dataset.… …They are used for testing OCR models that are part of the HuggingFace Transformers library. See here for details.…
Borealis + Agri-environmental Research Data Dataverse2022 · dataset · unknown
Fully Annotated Gas Prices of America Dataset for Multi-Metric Extraction in the WildThe fully annotated Gas Prices of America Dataset for Multi-Metric Extraction (GPA4MME) is a large-scale dataset of metric-specific annotations for gas prices and gas signs obtained from Google Streetview throughout the 49 mainland United States of America. The annotations greatly expand the utility of the original GPA dataset for multi-digit, multi-number price detection and context mapping (deno