Hugging Face Datasets2026 · Image
U(1) Defect Phase-Crossover ExperimentU(1) defect phase-crossover experiment Numerical study of the smallest eigenvalue of the unnormalized scalar connection Laplacian of an n×n open square grid with exactly one phased edge (plus 2D torus and 3D box extensions). Part of a multi-agent project on local frustration vs global spectral visibility. Start here: REPORT.md (v3, post-audit) — results with [T]/[A]/[N]/[C]/[O] evidence labels and
Hugging Face Datasets2026 · Image
RawVLA-BenchRawVLA-Bench RawVLA-Bench is a paired training-trajectory dataset for studying RAW-domain visual frontends for vision-language-action (VLA) policies. It is derived from successful expert trajectories in LIBERO and RoboTwin 2.0. Each trajectory has a RAW input version captured or rendered under a sampled lighting condition and a paired default-light RGB target. The pair shares the same trajectory i
Hugging Face Datasets2026 · Image
huanglianai/blogHugging Face Datasets2026 · Image
DOCO ImageNet-L and OOD-L with LAION-C DistortionsDOCO ImageNet-L and OOD-L with LAION-C Distortions This repository contains the materialized ImageNet-L and OOD-L image sets used for the LAION-C experiments in Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensation (DOCO) (CVF PDF, arXiv:2604.21772, code). The artifacts were generated by applying a modified version of the LAION-C corruption pipeline to ImageNet validation
Hugging Face Datasets2026 · Image
LeRobot + FiftyOne Blog FiguresLeRobot + FiftyOne Blog Figures Figures for the blog post "50 Embodiments, 497 Episodes, One Indexed Dataset: LeRobot + FiftyOne". The post walks through pulling 10 episodes from every robot embodiment in lerobot/community_dataset_v3, importing them into FiftyOne, embedding each episode with Qwen3-VL-Embedding-2B, and curating with similarity, uniqueness, and representativeness. The dataset the fi
Hugging Face Datasets2026 · Image
Holo4 TrajectoriesHolo4 Trajectories Every run behind the Holo4 benchmark scores: 7,366 agent trajectories from Holo4 27B and Holo4 35B-A3B, with each step's reasoning, actions, tool results and screenshots. Browse them at trajectories.hcompany.ai. This dataset is the same bundle the site serves. Benchmark Holo4 27B Holo4 35B-A3B Upstream License OSWorld 1,096 1,102 xlang-ai/OSWorld Apache 2.0 OSWorld 2 106 106 xla
Hugging Face Datasets2026 · Image
datachain/BVD-V-55M-CCBVD-V-55M-CC — the Creative Commons subset Every video in laion/BVD-V-55M-URLs whose YouTube licence is creativeCommon, in the same format as the original. The licence is not part of BVD, so it was resolved for the whole corpus through the YouTube Data API: all 2,119,224 source videos, one videos.list(part=status) call per fifty ids. 15,658 came back creativeCommon, 0.74%. This is a complete censu
Hugging Face Datasets2026 · Image
DuckerMaster/Thai-Synth-ReceiptsThai-Synth-Receipts Thai-Synth-Receipts is a large-scale, highly robust synthetic dataset of Thai commercial documents designed specifically for training and evaluating state-of-the-art Document AI and Optical Character Recognition (OCR) models. The dataset consists of 14,976 high-resolution document images (Receipts, Tax Invoices, Thermal Slips, and Quotations) across three distinct degradation v
Hugging Face Datasets2026 · Image
OmniTaskonomy Recipe DataOmniTaskonomy Recipe Data Paired image-to-image (I2I) and image-to-text (I2T) tasks for the R1–R6 training recipes and gradient analysis in OmniTaskonomy. Each of the six subsets has train and val splits. One row contains both objectives for the same task instance. from datasets import load_dataset data = load_dataset("Wakals/OmniTaskonomy_Recipe_Data", "jigsaw", split="train", streaming=True) sam
Hugging Face Datasets2026 · Image
MiniMax-H3 video benchmark mediaMiniMax-H3 video benchmark media Reference images, videos and audio for the MiniMax-H3 video benchmark, hosted so an inference server can fetch them by URL: https://huggingface.co/datasets/zhenghaoniTT/minimax-h3-bench-assets/resolve/main/<file> Attribution Videos, audio and the Tears of Steel stills are derived (trimmed, cropped, re-encoded) from Tears of Steel, (c) copyright Blender Foundation |
Hugging Face Datasets2026 · Image
EmbRACEEmbRACE: Embodied Reasoning and Action in Complex Environments Viewer · Paper · Code EmbRACE consists of two parts. The dataset is a set of 3,421 human demonstrations for closed-loop tasks in 55 environments, 48,264 steps in all, in which every step is paired with a rationale written from the agent's viewpoint and verified by annotators. The benchmark is a set of 686 tasks in 7 further environment
Hugging Face Datasets2026 · Image
OneJev-DataThe training data of OneJev: 94,707 typed questions about screens, photos, videos and text, each with its answer. This is 95.5% of the rows OneJev was trained on; rows whose sources do not allow redistribution are left out. Use from datasets import load_dataset ds = load_dataset("OmniJev/OneJev-Data", split="train", streaming=True) row = next(iter(ds)) Each row has a state with <image:N> and <vide
Hugging Face Datasets2026 · Image
OmniTaskonomyOmniTaskonomy OmniTaskonomy groups visual tasks into Recognition, Reconstruction, and Reorganization. Each task and modality has its own split, named family__task__i2i or family__task__i2t. The i2t and i2i configs group the evaluation and training splits, respectively. This release contains 9,444 I2T evaluation samples across 25 tasks and 350,000 I2I training samples across 7 tasks. I2T rows have
Hugging Face Datasets2026 · Image
VidScribeVidScribe VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12). Release status. This repository hosts the public hal
Hugging Face Datasets2026 · Image
VGDL-fMRI — Reason to PlayVGDL-fMRI: Reason to Play Human video-game learning, fMRI recordings, model gameplay and representations for Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners, accepted at NeurIPS 2026. Research code · Interactive results · Original human dataset TL;DR: Explore the replays on the website, or download the human recordings, model features and processed fMRI
Hugging Face Datasets2026 · Image
VietTravelVQA v2VietTravelVQA v2 VietTravelVQA v2 is a Vietnamese visual question answering dataset about tourism and cultural heritage in Vietnam. This release contains 9,530 question-answer pairs associated with 1,406 images. It combines the original 7,030 annotated pairs with 2,500 additional knowledge-grounded pairs. Dataset summary Split Question-answer pairs Images Train 6,805 1,051 Validation 1,010 194 Tes
Hugging Face Datasets2026 · Image
荆楚文化文物语义分析样本荆楚文化文物语义分析样本 这是用于审核字段设计和语义抽取质量的样本版本,共 116 条记录、35 个核心字段。 数据集另含 enrichment_trial_5 配置:从主表选取 5 条文物进行检索、图像观察与语义补充,共 40 个字段。主表内容未被覆盖。 湖北省博物馆:83 条 荆州博物馆:33 条 图片位于第 2 字段 image_url 两个古籍书影汇总页已拆分为 18 条单书记录 删除了当前来源完全无法填充的 creation_place 和 collection_number 删除派生检索字段 keywords;删除与 archaeological_site 高度重复的 provenance 原文与语义归纳分离;缺少来源的信息保持为空 所有记录目前均为 待人工复核,尚不是最终 604 条全量版本 数据文件 data/artifacts.csv:Dataset Viewer 使用的
Hugging Face Datasets2026 · Image
dm-bench 0.1.0dm-bench 0.1.0 Torn-document reassembly benchmark. Synthetic, seeded, licence-clean pages are torn into non-overlapping fragments; a solver must place every fragment back on the page with a rigid pose. Full benchmark card, metrics and baseline: docs/BENCHMARK.md. Code: Arittra-Bag/Dataset-Maker. Contents tier val pages test-dev pages test pages easy 39 13 48 medium 37 16 50 hard 33 19 52 puzzles/<
Hugging Face Datasets2026 · Image
Bad Apple!! in Qwen3 attention: the text is the videoBad Apple!! in Qwen3 attention: the text is the video Every frame of Bad Apple!! is one line of plain text. Feed the line to an unmodified Qwen3, and the first layer's attention logits draw the frame. No weights are trained or changed, and there is no hidden channel: the picture comes only from which real words stand where. https://www.youtube.com/watch?v=eFAwXZe_fZI The same second (1:00–1:10), d
Hugging Face Datasets2026 · Image
LME-BenchLME-Bench Long-horizon Multi-turn image Editing benchmark, introduced in MT-OPSD: On-Policy Self-Distillation for Multi-Turn Image Editing. Existing multi-turn editing benchmarks stop at five turns. LME-Bench has 100 ten-turn sessions in which every instruction is applied to the previous turn's output, to measure whether an editor keeps following instructions and keeps its images intact over long
Hugging Face Datasets2026 · Image
Jing2237/LongGameReasoningLongGameReasoning Long-horizon game-playing trajectories with image observations, structured decisions, and tool calls. This repository combines two validated English SFT releases into one dataset. The training splits are assigned by whole game, with no game shared between splits. Layout sft/train.jsonl: 10,396 training records from 17 games. sft/dev.jsonl: 1,813 validation records from 6 games. s
Hugging Face Datasets2026 · Image
Missing-texter/newartsdatasetsHugging Face Datasets2026 · Image
Danish PersonasDanish Personas Danish Personas contains 100,000 synthetic adult personas designed to reflect Danish society at the aggregate level. The demographic foundation is sampled and calibrated from official public statistics published by Statistics Denmark (Danmarks Statistik, DST), including population, education, labour-market, origin, and job-function distributions. The primary use case is to provide
Hugging Face Datasets2026 · Image
virajwirle/Plant-disease-datasetsHugging Face Datasets2026 · Image
PolyTopoBenchPolyTopoBench PolyTopoBench is a benchmark for complex vector polygon generation from remote-sensing imagery. It evaluates whether a model can produce complete polygon topology, including interior rings (holes), rather than only exterior boundaries. Paper: PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery, NeurIPS 2026 (Evaluations and Datasets Track) Cod