Hugging Face Datasets2026 · dataset
Ali Data — Urdu & Roman Urdu CorpusAli Data — Urdu & Roman Urdu Corpus A large public-domain/open collection of Urdu (اردو) and Roman Urdu text for training Urdu language models. Stats 21,846,033 unique rows (deduplicated by exact text) Format: JSON Lines — each line: {"text": "...", "source": "..."} Scripts: Urdu (Arabic script) + Roman Urdu (Latin script) Sources Source Rows (approx) Hugging Face open datasets (news, QA, sentimen
Hugging Face Datasets2026 · dataset
Chronicling America (US Library of Congress) Historical NewspapersChronicling America (US Library of Congress) - Parquet Dataset A high-performance, columnar Apache Parquet dataset containing digitized, OCR-extracted historical American newspapers from the US Library of Congress Chronicling America / National Digital Newspaper Program (NDNP). Produced by streaming and transmuting massive Library of Congress preservation archives (.tar.bz2, METS/MODS, and ALTO XM
Hugging Face Datasets2026 · Table · Parquet
Global-Ocean-Science-Corpus🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned) A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institute
Hugging Face Datasets2026 · Table · Parquet
SINAI/ALIA-es-cultural-heritageDataset Introduction The ALIA Spanish Cultural and Heritage Corpus is a strategic open data infrastructure designed to support research and innovation in digital humanities, cultural analytics, and Spanish-language NLP. It consolidates heterogeneous official and academic repositories into a single curated dataset, enabling broad and structured access to cultural heritage documentation from Spain.
Hugging Face Datasets2026 · Text · gated
The Universal Sikh Research Library (1B+ Words)🪯 SikhLibrary A Universal, Comprehensive Research Corpus for the Sikh Community ਵਾਹਿਗੁਰੂ ਜੀ ਕਾ ਖਾਲਸਾ ਵਾਹਿਗੁਰੂ ਜੀ ਕੀ ਫ਼ਤਿਹ ਜੀ Production home: sikhi.io · Planned backend for: sikharchive.net 📖 Overview SikhLibrary centralizes Sikh scripture, exegesis, history, and scholarship — scanned, OCR'd, and machine-readable — into a single, standardized research corpus. It exists so that Sikh sch
Hugging Face Datasets2026 · Table · Parquet
ClimbMix 40B AzerbaijaniClimbMix 40B — Azerbaijani A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate. Dataset Summary This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning lar
Hugging Face Datasets2026 · dataset
REVE Pretrain (Open Subset)REVE Pretraining Dataset: Open Subset This repository contains the open-access subset of the data used to train REVE (Representation Learning for EEG). While the full pretraining corpus spans a wider array of private or restricted sources, this subset includes all recordings released under permissive licenses that allow for redistribution and open research. Dataset Details Description This subset
Hugging Face Datasets2026 · Table · CSV
ChameleonHugging Face Datasets2026 · Text
DatarrX/Burmese-English-Code-Mixed-Corpus🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱ A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research. Dataset Details Organization: DatarrX Creator: Khant Sint Heinn (Kalix Louis) Number of Rows: 1,111 Language: Burmese (Unicode) & English Mix Dataset Format: .txt Licens
Hugging Face Datasets2025 · Text
Burmese Text Corpus for Natural Language ProcessingBurmese Text Corpus For Natural Language Processing 🎫 Choose your language: 🌏 English Version | 🇲🇲 မြန်မာဗားရှင်း 🌏 English Version This dataset is a specifically curated text corpus for the Burmese language. It is intended to support Natural Language Processing (NLP) tasks, language model training, and research related to the Burmese language. 1. About the Dataset The primary goal of creating thi
Hugging Face Datasets2025 · Table · Parquet · gated
openlegaldata/court-decisions-germanyOpen Legal Data: Court Decisions Germany This dataset is a preprocessed version of an Open Legal Data data dump, spefically it contains German court decisions. The dataset was automatically generated and uploaded to the HF hub using oldp-toolkit. Available dumps Date Configs 2026-05-20 dump-20260520, dump-20260520-10k, dump-20260520-1k 2022-10-18 dump-20221018, dump-20221018-10k, dump-20221018-1k
Hugging Face Datasets2024 · Text
proj-persona/PersonaHubScaling Synthetic Data Creation with 1,000,000,000 Personas This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas: We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA
Hugging Face Datasets2024 · Table · CSV
mou3az/Question-Answering-Generation-ChoicesThe dataset is a merged compilation of QuAIL, RACE, and Cosmos QA datasets, having undergone preprocessing.
Hugging Face Datasets2024 · Table · Parquet
UzBookv2Dataset Card for UzBooks V2 Dataset Summary UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits: Split Description Examples lat Fully Latin-transliterated version 38,339 cyr Fully Cyrillic-transliterated version 38,339 What's New in V2? OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR Cleaner Text: Goog
Hugging Face Datasets2024 · Table · Parquet
Gazzetta UfficialeGazzetta Ufficiale 👩🏻⚖️⚖️🏛️📜🇮🇹 La Gazzetta Ufficiale della Repubblica Italiana, quale fonte ufficiale di conoscenza delle norme in vigore in Italia e strumento di diffusione, informazione e ufficializzazione di testi legislativi, atti pubblici e privati, è edita dall’Istituto Poligrafico e Zecca dello Stato e pubblicata in collaborazione con il Ministero della Giustizia, il quale provvede alla di
Hugging Face Datasets2023 · Table · Parquet
H-novel-corpusUpdate 12/07/2024: convert to parquet to download easier. Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make. (͡ ° ͜ʖ ͡ °) tags: socks, garter belt, foot fetish, ntr, netori..... Thanks Moleys/Numeron for the dataset donation.
Hugging Face Datasets2023 · Table · Parquet
UzBooksDataset Card for BookCorpus Dataset Summary In an effort to democratize research on low-resource languages, we release UzBooks dataset, a cleaned book corpus consisting of nearly 40000 books in Uzbek Language divided into two branches: "original" and "lat," representing the OCRed (Latin and Cyrillic) and fully Latin versions of the texts, respectively. Please refer to our blogpost and paper (Comin
Hugging Face Datasets2023 · Text
Open Australian Legal CorpusOpen Australian Legal Corpus ⚖️ The Open Australian Legal Corpus by Isaacus, a foundational legal AI research company, is the first and only multijurisdictional open corpus of Australian legislative and judicial documents. Comprised of 229,122 texts totalling over 60 million lines and 1.4 billion tokens, the Corpus includes every in force statute and regulation in the Commonwealth, New South Wale
Hugging Face Datasets2023 · Text
NoticiasGovBrHugging Face Datasets2023 · Table · CSV
ChatGPT Jailbreak PromptsDataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English]
Hugging Face Datasets2023 · Table · Parquet · gated
Chilean Spanish CorpusChilean Spanish Corpus Descripción del dataset Este corpus se compone de textos en español de Chile en diversos contextos y plataformas, incluyendo: Noticias Contenido web Reclamos Tweets Fuentes La mayoría del contenido proviene de dos fuentes principales: Sitios web con dominio .cl: Extraídos del conjunto de datos mc4 corp, disponible en Hugging Face. Contenido scrapeado: Recopilado directamente
Hugging Face Datasets2023 · dataset · gated
NACHOSIn this paper, we propose an original study of PLMs in the medical domain on French language. We compare, for the first time, the performance of PLMs trained on both public data from the web and private data from healthcare establishments. We also evaluate different learning strategies on a set of biomedical tasks. In particular, we show that we can take advantage of already existing biomedical PL
Hugging Face Datasets2023 · Table · Parquet · gated
SpanishBooksSpanish Books Dataset Summary Dataset of books in Spanish crawled from web and torrents. Preprocessing Preprocessing performed by spanish_nlp. Licensing Information The dataset is available under the Creative Commons Attribution-ShareAlike License (CC BY-SA 4.0). Some books may be subject to copyright. Use for academic purposes only. Citation Information @misc{ortiz2022esbooks, title={Crawled Span
Hugging Face Datasets2023 · Text
US public firm Annual Reports (10-K)The dataset contains the annual report of US public firms filing with the SEC EDGAR system. Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences. Sentiment labels are provided on a per filing basis from the market reaction around the filing data. Additional metadata for each filing is included in the dataset.
Hugging Face Datasets2022 · Table · Parquet
Norod78/HebrewStageAndLyricsWithNewLinesDataset Card for "HebrewStageAndLyricsWithNewLines" Contains poems and stories from "New Stage" ("במה חדשה") Contains text lines from various Hebrew song lyrics Data contains new-line characters Generated from a text file in which different poems were seperated using a double new-line character The script I made for converting the text file into a dataset is available here