Hugging Face Datasets2026 · Table · Parquet
OpenJevData-140kDataset Card for OpenJevData-140k Dataset Summary OpenJevData-140k is a curated release of the data collection used to train OpenJev-4B. It contains 146,738 decision-making examples across 19 task categories, organized into SFT and RL splits. Each example presents a state, a question, and a request-specific set of natural-language options. The data include hard answers and soft probability distrib
Hugging Face Datasets2026 · Table · Parquet
pigProfessional/so-arm101-stack-green_20260929_160851This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/pigProfessional/so-arm10
Hugging Face Datasets2026 · Table · Parquet
Typed Decisions (Japanese)Typed Decisions — 日本語版(typed-decisions-ja v3) LocalLLaMA/typed-decisions の日本語版です。1 つの state(業務の状況)に対して型付きの質問をまとめて答え、それぞれの答えを確率分布で返す課題です。 数値は、すべて実際に測った値です。測っていないものは、そう書いています。 概要 元データ: LocalLLaMA/typed-decisions、revision f7a2487edd7a043a5441a5e9ccc7fe5ddbd9ebe8(Apache-2.0)。 翻訳器: Qwen3.5-35B-A3B(Ollama の qwen3.5:35b-a3b-q4_K_M、digest 3460ffeede5453ead027dbd2f821b12ad0aa3de54630971993babdb2165221f7、Ap
Hugging Face Datasets2026 · Table · Parquet
claimcheck-bench: Agent Success-Claim Verificationclaimcheck-bench A synthetic benchmark for checking whether an AI agent's success claim is supported by its tool results and environment state. An agent can say “done” after a failed write, an action on the wrong record, or an operation that never persisted. This dataset contains 300 labelled agent traces for evaluating detectors that distinguish successful completion from false success claims acr
Hugging Face Datasets2026 · Table · Parquet
EvalSafe O*NETEvalSafe O*NET 150 documents · 7,500 consensus-labeled questions · 9 candidate models. Snapshot: 2026-09-29. Default reference: consensus. Only questions with an available consensus target and their corresponding documents and final model results are included. The documents are synthetic workplace examples. The reference targets are model-generated, using Astra (gpt-6-astra) and Fable (claude-fabl
Hugging Face Datasets2026 · Table · Parquet
EinzzCookie/TikTok-Video-User-DataTikTok Reposts & Authors Dataset This dataset contains archived TikTok reposts and user/author metadata scraped continuously via public TikTok API endpoints. It is split into two Parquet files for easy relational querying. 📄 Dataset Structure The dataset consists of two Parquet files: 1. videos.parquet Contains metadata for individual reposted TikTok videos. Column Type Description id String Uniqu
Hugging Face Datasets2026 · Table · Parquet
ARGUSAt a glance · Quick start · Fields · Reading the labels · Citation ARGUS: Evidence-Grounded Auditing of Identification Assumptions Data release for "Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations" (Zhang, Xie, Parra & Correia), to appear at ClimateNLP 2026 (EMNLP 2026 workshop). ARGUS is a structured language-model pipeline that audits the evidence a
Hugging Face Datasets2026 · Table · Parquet · gated
krsKRS scrape every entry in Poland's National Court Register (Krajowy Rejestr Sądowy), about 1.2 million companies and associations, parsed from the full extract PDFs and updated weekly. from datasets import load_dataset krs = load_dataset("tiagozip/krs", split="train") format one parquet file, data/krs.parquet, one row per KRS number. column type notes krs string 10 digit KRS number register string
Hugging Face Datasets2026 · Table · Parquet
protodotdesign/magicbox-v1MagicBox v1 693,376 records for extraction, choice, binary classification (noul), and ordinal scoring. Format: minifield.magicbox/1.0. request_json contains the source text and field definitions. targets_json contains supervision. Parse these columns with json.loads. Source attribution, revisions, licenses, and transformations are recorded in sources.json. Dataset counts are recorded in report.jso
NCBI GEO2026 · Raster · GeoTIFF · 3 samples
A molecular and spinal circuit basis for the functional segregation of itch and painThe dorsal horn is remarkably diverse, yet how this cellular complexity enables discrimination of sensory modalities remains a fundamental question. Neurons expressing gastrin-releasing peptide receptor (Grpr+) have long been considered dedicated to itch, yet broad activation also evokes pain-related behavior. Whether Grpr+ neurons are required for pain, and how itch- and pain-related functions ar
Hugging Face Datasets2026 · Image
datachain/BVD-V-55M-CCBVD-V-55M-CC — the Creative Commons subset Every video in laion/BVD-V-55M-URLs whose YouTube licence is creativeCommon, in the same format as the original. The licence is not part of BVD, so it was resolved for the whole corpus through the YouTube Data API: all 2,119,224 source videos, one videos.list(part=status) call per fifty ids. 15,658 came back creativeCommon, 0.74%. This is a complete censu
Hugging Face Datasets2026 · Table · Parquet
Qwen3.8-Max DistillationQwen3.8-Max Distillation A quality-filtered derivative of Qwen3.8-Max Distillation 50K by r0b0tlab, prepared for local training and fine-tuning on consumer hardware. This repository takes the original 49,772-example dataset and produces a substantially smaller training set focused on coding, reasoning, instruction following, and tool use. [!CAUTION] Terms and provenance notice — not cleared for un
Hugging Face Datasets2026 · Table · Parquet
BBuckz/basketball-encyclopediaHugging Face Datasets2026 · Table · Parquet
typesafe/evalsafe-invoice-processingInvoice processing Snapshot: 2026-09-28. 150 cases and 6,874 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by po
Hugging Face Datasets2026 · Image
OmniTaskonomy Recipe DataOmniTaskonomy Recipe Data Paired image-to-image (I2I) and image-to-text (I2T) tasks for the R1–R6 training recipes and gradient analysis in OmniTaskonomy. Each of the six subsets has train and val splits. One row contains both objectives for the same task instance. from datasets import load_dataset data = load_dataset("Wakals/OmniTaskonomy_Recipe_Data", "jigsaw", split="train", streaming=True) sam
Hugging Face Datasets2026 · Table · Parquet
dhgottesman/paq-kas-clozeHugging Face Datasets2026 · Table · Parquet
PriceLab Product PricesPriceLab Product Prices I prepared this dataset for PriceLab, my product-price estimation project. It contains 19,991 training examples, 999 validation examples and 1,000 test examples derived from ed-donner/items_lite. I preserved the original split assignments and filtered descriptions containing explicit price expressions before training. What this dataset is for I use product descriptions to t
Hugging Face Datasets2026 · Table · Parquet
haijian06/yellow_cube_brown_box_v1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/haijian06/yellow_cube_br
Hugging Face Datasets2026 · Image
OneJev-DataThe training data of OneJev: 94,707 typed questions about screens, photos, videos and text, each with its answer. This is 95.5% of the rows OneJev was trained on; rows whose sources do not allow redistribution are left out. Use from datasets import load_dataset ds = load_dataset("OmniJev/OneJev-Data", split="train", streaming=True) row = next(iter(ds)) Each row has a state with <image:N> and <vide
Hugging Face Datasets2026 · Image
OmniTaskonomyOmniTaskonomy OmniTaskonomy groups visual tasks into Recognition, Reconstruction, and Reorganization. Each task and modality has its own split, named family__task__i2i or family__task__i2t. The i2t and i2i configs group the evaluation and training splits, respectively. This release contains 9,444 I2T evaluation samples across 25 tasks and 350,000 I2I training samples across 7 tasks. I2T rows have
Hugging Face Datasets2026 · Table · Parquet
Whittle distillation set (Qwen3.8-27B reasoning traces + top-20 logprobs)Whittle distillation set: Qwen3.8-27B traces with top-20 logprobs About 25,000 complete, quality-filtered answers from Qwen3.8-27B (thinking on), each with the teacher's top-20 log-probabilities at every answer token. It is built for logit-level knowledge distillation: a student can match the teacher's full per-token distribution, not just its sampled text. We use it to train Whittle-Qwen-3.8-35B-
Hugging Face Datasets2026 · Image
VidScribeVidScribe VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12). Release status. This repository hosts the public hal
Hugging Face Datasets2026 · Image
VGDL-fMRI — Reason to PlayVGDL-fMRI: Reason to Play Human video-game learning, fMRI recordings, model gameplay and representations for Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners, accepted at NeurIPS 2026. Research code · Interactive results · Original human dataset TL;DR: Explore the replays on the website, or download the human recordings, model features and processed fMRI
Hugging Face Datasets2026 · Image
VietTravelVQA v2VietTravelVQA v2 VietTravelVQA v2 is a Vietnamese visual question answering dataset about tourism and cultural heritage in Vietnam. This release contains 9,530 question-answer pairs associated with 1,406 images. It combines the original 7,030 annotated pairs with 2,500 additional knowledge-grounded pairs. Dataset summary Split Question-answer pairs Images Train 6,805 1,051 Validation 1,010 194 Tes
Hugging Face Datasets2026 · Table · Parquet
Aaaay-0610/so101-task-1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Aaaay-0610/so101-task-1.