Hugging Face Datasets2026 · Table · CSV
Old Games Transcript90s Games Transcript A small English-language corpus of narrative text from classic PC games. This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading. Dataset Config Records Content caesar3 20 Mission briefings + victory messages diablo2_lod 7 Cinematic
Hugging Face Datasets2026 · dataset
TUSI-BenchTUSI-Bench TUSI-Bench is a Persian legal benchmark designed for research on legal question answering, legal knowledge retrieval, and evaluation of large language models in the domain of Iranian law. TUSI-Bench consists of three complementary components: TUSI-KB: a structured Persian legal knowledge base containing legal provisions from Iranian law. TUSI-QA: an open-ended Persian legal question-ans
Hugging Face Datasets2026 · Text
SPIRAL-Bench v0SPIRAL-Bench v0 A small benchmark for testing whether the wording of a retrieval query changes the balance of the evidence a retriever returns. Built for the Ouroboros project, which studies self-confirming retrieval loops in agentic RAG. The question this dataset exists to answer In agentic RAG, the system writes its own follow-up search queries, and it writes them using what it currently believe
Hugging Face Datasets2026 · Text
OEP-BenchOEP-Bench OEP-Bench (Orex Enterprise Policy Benchmark) is a synthetic benchmark for evaluating retrieval-augmented generation (RAG) and enterprise knowledge systems under policy versioning, access-control, citation, temporal-reasoning, and security constraints. OEP-Bench is built around NovaCore Technologies, a fictional enterprise used solely as the organizational setting for the benchmark. NovaC
Hugging Face Datasets2026 · dataset
Chronicling America (US Library of Congress) Historical NewspapersChronicling America (US Library of Congress) - Parquet Dataset A high-performance, columnar Apache Parquet dataset containing digitized, OCR-extracted historical American newspapers from the US Library of Congress Chronicling America / National Digital Newspaper Program (NDNP). Produced by streaming and transmuting massive Library of Congress preservation archives (.tar.bz2, METS/MODS, and ALTO XM
Hugging Face Datasets2026 · Table · Parquet · gated
medical-bench-TWDataset Card for medical-bench-tw medical-bench-tw 是以台灣公開官方來源策展的繁體中文醫療/照護評測資料集,涵蓋考選部醫事人員國考、專科/次專科醫師筆試、專科護理師、入學測驗(會考/學測)選擇題,以及健保/食藥/疾管/學會等文件語料。公開 HQ 版提供約 17 萬 題單選(整卷 train/validation/test)與進階複選軌,並以多個 Hub config(大包+細拆)方便載入。 卡片結構參考 lianghsun/tw-emergency-medicine-bench(Formosa-bench 系急診專科評測卡);本資料集範圍更廣、schema 為本專案自訂欄位(非整份 Formosa-bench 扁平 A–E 格式)。 怎麼看分類? Hub 頁面的 Subset / Config 下拉 = 下方各表的 config 欄。下拉名
Hugging Face Datasets2026 · Table · Parquet
EMNLP 2026 Skill Retrieval DatasetsWhen Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents This Readme file serves as a guide for using the two datasets we presented in our paper: Harbor Trial dataset which consists of real task execution data and Track B dataset which contains synthetic data. Those datasets are intended to finetune your own skill retrieval solutions and recipes and serve as a common
Hugging Face Datasets2026 · Table · Parquet
arXiv Complete CorpusarXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions
Hugging Face Datasets2026 · Text
The Bible in 1,004 LanguagesThe Bible in 1,004 Languages 14,497,397 verses across 1,253 translations in 1,004 languages, every verse keyed to the same chapter-and-verse address so that any two languages can be aligned by joining on book, chapter and verse. The Bible is the most widely translated text in existence, and for several hundred of the languages here it is the largest — sometimes the only — substantial digitised tex
Hugging Face Datasets2026 · Text
Massachusetts Case LawMassachusetts Case Law Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Source & credit — Free Law Project / CourtListener Every opinion in this dataset comes from the Free Law Project / CourtLis
Hugging Face Datasets2026 · Text
pablobenavidesj/doctrina-jurisprudencia-chile🇨🇱 Corpus Jurídico y Doctrinal de Chile en Markdown (Open Legal Chile) Bienvenido al repositorio oficial del Corpus Jurídico Canónico, Doctrinal y Jurisprudencial de Chile, desarrollado y mantenido por Open Legal Chile. Este repositorio ofrece acceso 100% completo, libre y gratuito (Apache-2.0) al texto íntegro de la dogmática jurídica chilena, a las Guías Oficiales de la Academia Judicial, a los
Hugging Face Datasets2026 · Table · Parquet
Italy Legislation (Normattiva OpenData)Italy Legislation (Normattiva OpenData) Research snapshot of official national legislation from Normattiva OpenData (Poligrafico) official BFF. Not legal advice. The official gazette / authentic source prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-07 Coverage catalog-backed incomplete (203,112 fetched / 204,668 discovered; 1,556 residual failures) Source Normattiva OpenData
Hugging Face Datasets2026 · Table · Parquet
Türk İçtihat Korpusu (11M Karar)Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9
Hugging Face Datasets2026 · dataset
RETECO SemEval-2027 Training and Development DataRETECO Training & Development Data · SemEval-2027 Task 1 Retrieval that must reason about when evidence applies and what the conversation has already established. 📋 Task · 📊 Data · 📐 Evaluation · 🚀 Participate · 🧰 Starter kit 🎯 At a glance What Official training and development data for RETECO, the SemEval-2027 shared task on reasoning-oriented retrieval Scope 2 tracks · 5 sub-tracks · 24 self-con
Hugging Face Datasets2026 · Table · Parquet
Türk İçtihat Korpusu (11M Karar)Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9
Hugging Face Datasets2026 · Table · Parquet · gated
Open India LawOpen India Law Open, structured Indian primary law - plus the scrapers that build it. Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15 tribunals and regulators, and Central, State and Union Territory legislation down to the individual section. Normalized to one schema, exclusively from official government sources. Volume Period Court judgments 12,848,644 195
Hugging Face Datasets2026 · Text
PerSHOP Response RankingPerSHOP Response Ranking A Persian response-ranking benchmark derived from PerSHOP and used to evaluate lexical, semantic-embedding and LLM-based ranking methods. It includes the existing Random and Same-Domain configurations alongside a newly developed Domain-Lexical hard-negative configuration. Each instance pairs a customer query with 5 candidate responses: 1 correct (gold) and 4 negatives. The
Hugging Face Datasets2026 · Table · Parquet
CodeGraphCodeGraph An open-taxonomy, Wikidata-grounded semantic knowledge graph over 145M source files. CodeGraph annotates 144,910,008 source files from Stack-Edu, across 14 programming languages, along four orthogonal semantic axes — application domains, algorithms (with category and asymptotic complexity), programming paradigms, and design patterns — and grounds the resulting concept vocabulary in Wikid
Hugging Face Datasets2026 · Image
Indonesian KTP Dataset 24K (Flat & 3D Perspective Augmented)Indonesian KTP Dataset 24K (Flat & Augmented - Commercially Safe) Welcome to the Indonesian KTP (Kartu Tanda Penduduk) Dataset. This is a highly robust, high-fidelity, and commercially safe synthetic dataset designed to advance SOTA (State-of-the-Art) research in Document Information Extraction (DIE), Key Information Extraction (KIE), and Optical Character Recognition (OCR) specifically for Indone
Hugging Face Datasets2026 · Table · Parquet · gated
Open US Law - Statutes, Regulations, Court Rules, Agency Guidance and DecisionsOpen US Law New in v2026.09 Updated September 29, 2026. 5,276,632 sections, up from 2,978,617 in v2026.08 (+2,298,015). 317 files, up from 229, including 95 new ones. Every section appears exactly once and carries the official URL it was published from. Regulations for 49 jurisdictions, alongside the current Code of Federal Regulations. 104,540 state attorney general opinions from 42 jurisdictions
Hugging Face Datasets2026 · dataset
Embodied AI Literature MetadataEmbodied AI Literature Metadata This dataset contains normalized paper metadata collected for an Embodied AI / Vision-Language-Action literature assistant. It is intended for metadata search, paper triage, and PDF retrieval before PaperQA-style evidence reading. Generated at: 2026-07-04T04:32:33.578443+00:00 Splits Split Records With abstract With PDF URL topconf_all 79068 24519 62689 frontier_202
Hugging Face Datasets2026 · Table · CSV
TheoremGraph MatchingTheoremGraph Matching Formal–informal theorem matches from the TheoremGraph paper. Each row pairs a Lean declaration with the most similar natural-language statement from arXiv, found by cosine similarity over slogan embeddings, and labeled by an LLM judge as exact, inexact, or wrong (the first two count as a match). The file contains every candidate pair at cosine similarity 0.80 and above: 100,8
Hugging Face Datasets2026 · Table · CSV
uw-math-ai/math-graphMath-Graph Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency graph spanning both informal and formal mathematics. On the informal side it parses millions of theorem-like environments from mathematics arXiv and recovers directed dependency edges within and across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed declaration
Hugging Face Datasets2026 · Text
IconClip Search BenchmarkIconClip Search Benchmark A BEIR-format information-retrieval benchmark for icon search by text intent. 120 hand-curated paraphrase queries against a 22 827-icon corpus spanning 11 open-license icon libraries (Lucide, Phosphor, Tabler, Heroicons, Bootstrap, Carbon, Font Awesome, Iconoir, Ionicons, Material Symbols, RemixIcon). The dataset supports two retrieval tasks on the same queries / qrels: T
Hugging Face Datasets2026 · Text
SkillRet BenchmarkSkillRet Benchmark 📄 Technical report: SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents (arXiv:2605.05726) Dataset Overview SkillRet is a retrieval benchmark for matching natural-language user requests to agent skills. It contains a curated library of public agent skills from GitHub with synthetic training and evaluation queries. Dataset Statistics Metric Value Total Records 218