Hugging Face Datasets2026 · Table · Parquet
Federal Reserve Beige BookFederal Reserve Beige Book Every edition of the Beige Book, the Federal Reserve's Summary of Commentary on Current Economic Conditions by Federal Reserve District, from the first, May 20, 1970, when the Board's pages call it the Redbook, to the latest. One row per section: the national summary, each of the twelve districts, and the one special report (May 18, 1983). In the words of the Board's Bei
Hugging Face Datasets2026 · Text
H2S-Research/Highlight-Then-SummarizeHighlight-Then-Summarize Highlight-Then-Summarize (H2S) is a compress-then-reason approach for long-context understanding. It makes evidence localization and information integration explicit before final-answer generation: long document + question -> evidence -> question-conditioned summary -> answer [Code] [Data format] H2S highlights source-addressable evidence from a block-structured document,
Hugging Face Datasets2026 · Table · Parquet
US Bill and Resolution Summaries (Congressional Research Service)US Bill and Resolution Summaries (Congressional Research Service) Every summary of a bill or resolution of the United States Congress that the Congressional Research Service (CRS) wrote and the Congress.gov API lists, from the 93rd Congress (1973-1974) on, with its text. CRS summarizes a measure when it is introduced and again at later actions, such as passing a chamber, so a bill can have several
Hugging Face Datasets2026 · Table · Parquet
RuEn Code CurriculumRuEn Code Curriculum RuEn Code Curriculum is a curated Russian-English dataset for continued pretraining (CPT) and supervised fine-tuning (SFT) of small code-oriented language models. This public release contains only records classified as redistributable. Local-training-only web and code sources used by the internal curriculum are intentionally excluded. Dataset summary Configuration Split Record
Hugging Face Datasets2026 · Table · Parquet
US Congressional Research Service ProductsUS Congressional Research Service Products Every product of the Congressional Research Service (CRS), the research service of the United States Congress, that the Congress.gov API lists, active and archived, with full text and metadata: Reports, Posts, Resources, Testimony and Infographics, as the API names them. CRS writes them for Members of Congress; Congress.gov publishes them. Nothing here is
Hugging Face Datasets2026 · Table · Parquet
Uzbek-law-LM: Authoritative Legal Corpus & SFT BenchmarkUzbek-law-LM: Authoritative Legal Corpus & SFT Benchmark Uzbek-law-LM is an engineered, legally grounded, and decontaminated Uzbek legal corpus designed for pre-training, continued pre-training, and supervised fine-tuning (SFT) of large language models serving the Republic of Uzbekistan's legislative, judicial, executive, and civic sectors. It provides a single, cohesive, authoritative legal resou
Hugging Face Datasets2026 · Table · Parquet
Türk İçtihat Korpusu (11M Karar)Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9
Hugging Face Datasets2026 · Table · Parquet
Türk İçtihat Korpusu (11M Karar)Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9
Hugging Face Datasets2026 · dataset
unohamza/Arabic-news-dailyArabic News Daily 🗞️ A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources. Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains. Sources Source Domain Variety Al Jazeera Arabic Politics MSA BBC Arabic Politics MSA RT Arabic Politics
Hugging Face Datasets2026 · Table · Parquet
TimeGround-1MTimeGround-1M Synthetic English audio dataset for time-aware speech understanding, covering temporal localization, temporal description, and timed summaries. Data Filtering We use 14k hours of audio from YODAS2 English shards, selected from a 24k-hour source pool after language- and silence-ratio filtering. Synthetic annotations were generated for three time-grounded tasks, then filtered through L
Hugging Face Datasets2026 · Text
BeTraC 2026 - DoPaCo Audio DatasetBeTraC 2026 - Synth-DoPaCo Audio Dataset Synthetic doctor-patient conversations with audio, transcripts, dialog metadata, and SOAP note summaries. Dataset Splits Split Dialogs Shards Size dev 400 1 469 MB train 7,200 9 8.7 GB File Format This dataset uses the WebDataset format (tar archives). Each sample contains 4 files sharing the same key (e.g., dialog_0060_0120): Extension Content .opus Opus-c
Hugging Face Datasets2026 · Table · Parquet
Polish Court DecisionsPolish Court Decisions The largest open dataset of Polish court decisions: 2,830,029 decisions with full texts across all court levels. What Makes This Dataset Unique Source This dataset Best on HF (JuDDGES) Difference Common courts 437,446 437,450 (pl-court-raw) same source Administrative courts 1,899,852 ~1,800,000 (pl-nsa) same source Supreme Court + Constitutional Tribunal + KIO 492,731 0 +493
Hugging Face Datasets2026 · Text · gated
CyberSecurity-1MCyberSecurity-1M A large-scale, multi-source cybersecurity knowledge dataset containing 1.19M records across 16 categories, collected exclusively for academic, non-commercial research purposes. Last updated: 2026-05-27. Disclaimer: This dataset is provided for academic research only. All content is aggregated from publicly available sources. The views, opinions, and information expressed in the da
Hugging Face Datasets2026 · Table · CSV
ChameleonHugging Face Datasets2026 · Table · Parquet
CapTrackDataset Card for CapTrack Dataset Summary CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary dimensions: CAN (Latent Competence): What a model is capable of doing under ideal prompting WILL (Default Behavioral Preferences): What a mod
Hugging Face Datasets2025 · Image
FLAWSFLAWS: Faults Localization Across Writing in Science FLAWS is a benchmark for evaluating error identification and localization in scientific papers. It currently consists of 713 paper–error examples, including: 265 unique papers with one error inserted using GPT-5 (in ALL_OPENAI.tar.gz) 448 unique papers with one error inserted using Gemini 2.5 Pro (in ALL_GEMINI.tar.gz) The dataset is generated u
Hugging Face Datasets2025 · Text
moltextnetLoad the MolTextNet Dataset from datasets import load_dataset # Load the dataset dataset = load_dataset("liuganghuggingface/moltextnet") # Print the dataset object print(dataset) # Output: # DatasetDict({ # train: Dataset({ # features: ['id', 'canonical_smiles', 'description'], # num_rows: 2474584 # }) # }) # Show available splits print(dataset.keys()) # Output: # dict_keys(['train']) # Display th
Hugging Face Datasets2025 · Table · CSV
leet-leetLeetCode Scraper Dataset This dataset contains information scraped from LeetCode, including problem details, metadata, and related files. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. questions_deets.csv Contains detailed information about each problem, including problem descriptions,
Hugging Face Datasets2025 · Table · Parquet
Smol-KoreanTalk: Korean SmolTalk DatasetSmolLM2의 인스트럭션 훈련 데이터 HuggingFaceTB/smol-smoltalk를 한국어로 번역했어요.
Hugging Face Datasets2025 · Table · CSV
Customer Support TicketsFeaturing Labeled Customer Emails and Support Responses 🔧 Synthetic IT Ticket Generator — Custom Dataset Create a dataset tailored to your own queues & priorities (no PII). 👉 Generate custom data Define your queues, priorities, language Need an on-prem AI to auto-classify tickets?→ Open Ticket AI There are 2 Versions of the dataset, the new version has more tickets, but only languages english and
Hugging Face Datasets2024 · dataset · gated
dongqi-me/VISTAVISTA Dataset Dataset Structure dataset/ ├── videos/ # Video files directory │ ├── train_part1/ # Training set videos (first 8000 samples) │ ├── train_part2/ # Training set videos (remaining samples) │ ├── val/ # Validation set videos │ └── test/ # Test set videos ├── train_part1.json # Training set metadata (first 8000 samples) ├── train_part2.json # Training set metadata (remaining samples)… See
Hugging Face Datasets2024 · Table · Parquet
ramadita/indo-islamic-articleIndo-Islamic Article Dataset The Indo-Islamic Article Dataset contains Islamic articles in Indonesian, collected from various web sources. The dataset supports tasks such as text classification, summarization, and sentence similarity. It was primarily used in the SEQURAN project, an Islamic knowledge-based search engine focusing on information retrieval and summarization. Columns: title: The title
Hugging Face Datasets2024 · Table · CSV
Greek WikipediaGreekWikipedia A Greek abstractive summarization dataset collected from the Greek part of Wikipedia, which contains 93,432 articles, their titles and summaries. This dataset has been used to train our best-performing model GreekWiki-umt5-base as part of our research paper:Giarelis, N., Mastrokostas, C., & Karacapilidis, N. (2024) Greek Wikipedia: A Study on Abstractive Summarization.For informatio
Hugging Face Datasets2024 · Text
FlashRAG Datasets⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your cus
Hugging Face Datasets2024 · Table · Parquet
RussianNLP/Mixed-Summarization-DatasetRussian summarization data mix Total Number of items in Train: 197561. Total Number of items in Golden Test set: 258 (manually verified semi-synthetic data for evaluation purpose). We use this datasets for train mix: XLSum Gazeta WikiLingua MLSUM Reviews (ru) Curation-corpus (ru) Matreshka DialogSum (ru) SAMSum (ru) Cite us @misc{akhmetgareeva2024summary, title={Towards Russian Summarization: can