{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:tabular-classification",
"evidence": null
}
],
"evidence_policy": "strict"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Table · Parquet
EinzzCookie/TikTok-Video-User-DataTikTok Reposts & Authors Dataset This dataset contains archived TikTok reposts and user/author metadata scraped continuously via public TikTok API endpoints. It is split into two Parquet files for easy relational querying. 📄 Dataset Structure The dataset consists of two Parquet files: 1. videos.parquet Contains metadata for individual reposted TikTok videos. Column Type Description id String Uniqu
Hugging Face Datasets2026 · Table · Parquet
AI Pioneer Program: Four-Year OutcomesAI Pioneer Program: Four-Year Outcomes, 2024–2027 MODELED DATA. Every record in this dataset is from a raw dataset representing a child in ZEN AI Co's AI Pioneer Program which is the first of it's kind and scale in United States history. The 55,000 student records were observed. Interactive explorer: ZENLLC/ai-pioneer-outcomes-explorer · Program: zenai.world/ailiteracyyouth · ZEN Arsenal: arsenal.
Hugging Face Datasets2026 · Table · Parquet
Would You Rather: Kissing vs Sauce (10k humans)Would You Rather: Kissing vs Sauce 💋🍝 [!NOTE] The analysis in this README was written by an AI (Claude, working as Rapidata's engineering assistant), and every number was computed by a script from the data in this repo. The data itself is real: 10,000 human responses collected through Rapidata. Treat the interpretation as exploratory, not peer-reviewed. We asked 10,000 people in 127 countries one
Hugging Face Datasets2026 · Image
The Food API: Packaged Food & Beverage Products Dataset (Sample)The Food API: Packaged Food & Beverage Products Dataset Normalized packaged-food and beverage product data, keyed by barcode, for food-tech, nutrition apps and e-commerce. This repository is the free evaluation sample of The Food API: one record per product with the barcode, the full ingredient statement and its parsed tree, declared and precautionary allergens, the nutrition panel, package claims
Hugging Face Datasets2026 · Image
The Food API: Packaged Food & Beverage Products Dataset (Sample)The Food API: Packaged Food & Beverage Products Dataset Normalized packaged-food and beverage product data, keyed by barcode, for food-tech, nutrition apps and e-commerce. This repository is the free evaluation sample of The Food API: one record per product with the barcode, the full ingredient statement and its parsed tree, declared and precautionary allergens, the nutrition panel, package claims
Hugging Face Datasets2026 · Table · CSV
Agent Memory Resilience & Poisoning BenchmarkAgent Memory Resilience & Poisoning Benchmark Dataset Summary This benchmark dataset evaluates resilience, negative transfer, and memory poisoning mitigation in autonomous LLM agent architectures (such as LangGraph, AutoGen, and CrewAI). When autonomous agents record distilled self-reflections after attempting tasks, external stochastic failures or subtle API deprecations often cause agents to com
Hugging Face Datasets2026 · Table · Parquet
Shingan FinRisk LabelsShingan FinRisk Labels The labelled dataset of the Shingan (心眼) project: risk labels and derived features for listed companies at specific as_of dates, built to train and evaluate a dual-track (structured GBDT + text QLoRA) risk model. This dataset contains neither raw news text nor the raw price panel. To train the text track you must obtain the underlying text yourself from the original sources
Hugging Face Datasets2026 · Table · CSV
davidev07/synthetic-ecommerce-datasetSynthetic E-commerce Dataset - Free Sample This repository contains a free sample of the Synthetic E-commerce Dataset: 100 users, 50 products, 1,000 transactions, with full relational integrity (foreign keys, temporal consistency, validated business rules) and a fixed random seed (42) for reproducibility. License below (CC0) applies only to this free sample, not to the full-size commercial packs.
Hugging Face Datasets2026 · Table · CSV
Football Charts — Match Results and Goal TimingFootball Charts — Match Results and Goal Timing Football Charts publishes results, fixtures, league tables and the minute of every goal for 93 leagues in 42 countries, including the lower divisions and women's competitions most sources skip. Free JSON API and MCP server; the results and goal-timing dataset is CC BY 4.0 with a DOI (10.5281/zenodo.22295583). Most public football datasets cover the b
Hugging Face Datasets2026 · Image · gated
OneAstronomy Strong Lens Master CatalogOneAstronomy Strong Lens Master Catalog (v1) Unified lens and non-lens tables compiled for the OneAstronomy strong-lensing workflow (2026).This release is a homogenized collection of public catalogs. Candidates found by our own search will be added later. The repository now includes catalog tables and Legacy Survey image cutouts (FITS + JPEG). Where the files are jady-zhao/OneAstro_StrongLens ├──
Hugging Face Datasets2026 · Table · CSV
chinna887/ecommerce-fraud-detection-synthetic-10k-sampl🛡️ Synthetic E-Commerce Fraud & AML Detection Dataset (10k Evaluation Sample) ⚠️ NOTICE: This is a truncated 10,000-row evaluation sample strictly for schema verification and local testing. 💳 [OBTAIN THE 10-MILLION ROW COMMERCIAL LICENSE HERE] > https://buy.stripe.com/8x26oIad4eH9eJf6gJ5wI01 🚀 Quick Start (Load via Hugging Face) Data scientists can instantly load this evaluation slice into their P
Hugging Face Datasets2026 · Table · CSV
SynSEPA — Synthetic SEPA Instant Payment Fraud DatasetSynSEPA: A Synthetic SEPA Instant Payment Dataset for APP Fraud Detection Research Overview SynSEPA is the first publicly available synthetic dataset specifically designed for Authorised Push Payment (APP) fraud detection in SEPA Instant Credit Transfer payments. No public SEPA fraud dataset previously existed. SynSEPA fills this gap by providing a large-scale, statistically calibrated synthetic d
Hugging Face Datasets2026 · Table · Parquet
hmda_2024HMDA 2024 (Home Mortgage Disclosure Act) Full-year 2024 loan application register (LAR) data released under the Home Mortgage Disclosure Act (HMDA), re-published here as a single Parquet file for convenient loading with the datasets library. Dataset summary Rows: 12,229,298 loan application records Columns: 99 (the full public LAR field set — property, applicant, underwriting, and pricing informat
Hugging Face Datasets2026 · dataset
EU Air Traffic LakeEU Air Traffic Lake Live and scheduled European air traffic as versioned Parquet: raw Bronze intake windows from the collector, plus curated Silver snapshots refreshed by a scheduled pipeline. Updates Bronze windows land continuously (positions every ~5 min, weather every 5 min, departures twice an hour, arrivals backfilled nightly). Each of the Silver snapshots below is replaced every 15 minutes
Hugging Face Datasets2026 · Table · Parquet
PumpFun Launch-to-Graduation CorpusPumpFun Launch-to-Graduation Corpus (Jun–Jul 2026) 798,430 pump.fun token launches. 33.58 million trades. 26.9 million bonding -curve snapshots. Every graduation outcome labeled. Tracked continuously, second by second, for 39 uninterrupted days. ⚠️ This dataset has documented, quantified data-quality issues — several are not optional to handle correctly. Full detail, root causes, and exact handlin
Hugging Face Datasets2026 · Table · Parquet
THBKG — Temporal Heterogeneous Biomedical Knowledge GraphTHBKG — Temporal Heterogeneous Biomedical Knowledge Graph A dated biomedical knowledge graph built from Open Targets 26.03 (with Reactome, ChEMBL and ClinicalTrials.gov), plus a clinical-advancement benchmark: rank target–disease pairs by their likelihood of advancing to Phase III, scored only from evidence datable strictly before each pair's decision year. Every temporal edge carries the year its
Hugging Face Datasets2026 · Text
US Parcel Layer — The Landrecords.us Nationwide Parcel DatasetUS Parcel Layer — The Landrecords.us Nationwide Parcel Dataset The Landrecords.us National Parcel Dataset is a comprehensive, standardized geospatial dataset aggregating ~157 million parcel boundaries and their associated land-ownership and taxation attributes, harmonized from thousands of local jurisdictions across the United States into a single national schema. Each record is a land parcel: a p
Hugging Face Datasets2026 · Table · Parquet
TimeSeventeen/Polymarket-v2Polymarket-v2 A large-scale dataset of on-chain event logs from Polymarket v2, the prediction-market platform on the Polygon network. The repository contains three layers covering the full contract lifecycle from Polymarket v2 server start: OrderFilled/ (the raw on-chain trade tape), daily_aligned/ (the cleaned, metadata-enriched, and normalized analysis layer, Split by UTC). update daliy.
Hugging Face Datasets2026 · Table · Parquet
Polymarket 5-Minute Crypto Up/Down MarketsPolymarket 5-Minute Crypto Up/Down Markets Second-by-second top-of-book order-book recordings of Polymarket's recurring 5-minute crypto up/down markets — bet on whether a coin closes a fixed 5-minute window higher or lower — for BTC, ETH, SOL, XRP, DOGE, HYPE and BNB. ~89,000 markets (one resolved 5-minute window each) 26,858,579 per-second observations (≈300 ticks per market) Span: BTC 24 Mar 202
Hugging Face Datasets2026 · dataset
adibmed/football-datasetFootball Predictions Dataset International football match dataset used to train an ML prediction system (CatBoost + Poisson/LightGBM ensemble) that predicts win/draw/loss probabilities and expected goals for any match between two national teams. Used to generate World Cup 2026 predictions — both pure ML and LLM consensus (Claude, GPT-5.5, Gemini, DeepSeek, Grok, Qwen). Dataset contents output/ — M
Hugging Face Datasets2026 · Table · Parquet
BeyondArena DatasetsBeyondArena Datasets Datasets from BeyondArena, a unified, holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types (with text, with high cardinality) from a broad range of disciplines. We introduce BeyondArena and its datasets in Beyond IID: How General Are Tabular Foundation Mod
Hugging Face Datasets2026 · Table · Parquet
Phishing Email Curated CleanedPhishing Email Curated Cleaned Phishing Email Curated Cleaned is a cleaned and AI-ready version of the original Phishing Email Curated Datasets by Champa, Rabbi and Zibran (2024), an aggregation of 11 heterogeneous email corpora released on Zenodo for benchmarking phishing email detection with machine learning. The original collection aggregates emails from public corpora spanning 1995–2022 (CEAS-
Hugging Face Datasets2026 · Table · CSV
AptaBenchAptaBench AptaBench is a benchmark for aptamer–small-molecule interaction prediction. It contains curated DNA/RNA aptamer–ligand pairs with standardized sequences, canonical SMILES, experimentally grounded active/inactive labels, quantitative affinity values where available, and fixed leakage-aware evaluation splits. This repository is provided for anonymous peer review. Author identities, affilia
Hugging Face Datasets2026 · Table · Parquet
MALTA-WLOGSMALTA-WLOGS Well-log feature tables extracted from the MALTA-WELD project. The default preview is the Features table, which combines conventional log measurements with lithology intervals. The other tables are available as separate dataset configurations. Configurations default: Features LogsFeatures: conventional well-log measurements AGPSummary: lithology summary by well AGPLithology: lithology
Hugging Face Datasets2026 · dataset
TopBenchTopBench Dataset TopBench is a benchmark for predictive reasoning over tabular data. Each example asks a model to infer an unobserved outcome, decision, treatment effect, or ranked/filtering result from historical tables and a natural-language query. Layout single_point_prediction/ decision_making/ treatment_effect_analysis/ ranking_and_filtering/ Each task directory contains dataset folders with