Hugging Face Datasets2026 · Table · CSV
Sintético de citações jurídicas (Desafio Jusbrasil BRACIS 2026)Sintético de citações jurídicas — Desafio Jusbrasil BRACIS 2026 Pareceres jurídicos sintéticos em português, com gabarito de cada citação (posição exata no texto e classe). Foram gerados pela equipe para medir e calibrar um verificador de citações. A regra do desafio exige que dado sintético usado em treino seja publicado, e este repositório cumpre isso. São duas versões, com o mesmo gabarito (mes
Hugging Face Datasets2022 · Image
Europeana NewspapersDataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiativ
Hugging Face Datasets2026 · Image
Ledger…Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values.…
Hugging Face Datasets2026 · Image
Synthetic Text Images (English)Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise.
Hugging Face Datasets2026 · Table · Parquet
artefactory/ledger-long-context-KPI-QALEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TRE
Hugging Face Datasets2026 · Image
VidScribeVidScribe VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12). Release status. This repository hosts the public hal