Hugging Face Datasets2026 · Image
datachain/BVD-V-55M-CCBVD-V-55M-CC — the Creative Commons subset Every video in laion/BVD-V-55M-URLs whose YouTube licence is creativeCommon, in the same format as the original. The licence is not part of BVD, so it was resolved for the whole corpus through the YouTube Data API: all 2,119,224 source videos, one videos.list(part=status) call per fifty ids. 15,658 came back creativeCommon, 0.74%. This is a complete censu
Hugging Face Datasets2026 · Image
VidScribeVidScribe VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12). Release status. This repository hosts the public hal
Hugging Face Datasets2026 · Image
Relative Camera Movement Benchmark (Image-to-Video)Rapidata Relative Camera Movement Benchmark Built by Rapidata. This dataset contains 900,318 human responses, collected with the Rapidata Python SDK, comparing how well 15 image-to-video models and world models move the camera relative to what is in the scene. Each row is a head-to-head comparison between two models' clips generated from the same still and the same instruction, judged by human ann
Hugging Face Datasets2026 · Table · Parquet · gated
Datapoint Text-to-Video Human Preferences (326K)Text-to-video human preferences: 326K votes across 15 models This dataset contains the complete voting record behind the Datapoint Video Bench leaderboard: 325,520 validated pairwise votes — exactly 10 for each of 32,552 video pairs. The votes compare 15 text-to-video models on 314 prompts built to stress motion, physics, and temporal consistency, judged by 22,982 annotators in 187 countries. Ever
Hugging Face Datasets2026 · Image
Camera Movement Benchmark (Image-to-Video)Rapidata Camera Movement Benchmark Built by Rapidata. This dataset contains 324,044 human responses, collected with the Rapidata Python SDK, comparing how well 15 image-to-video models and world models execute a described camera movement from a single still image. Each row is a head-to-head comparison between two models' clips generated from the same still and the same instruction, judged by human
Hugging Face Datasets2026 · Image
VideoArgusBenchVideoArgusBench Sample-specific rubric benchmark for conditioned video generation. VideoArgusBench is the evaluation benchmark for VideoArgus, a framework that scores a generated video against a rubric written for that specific prompt rather than a fixed global metric. This dataset ships the inputs (conditioning assets + prompts) and, for each input, a rubric. It does not contain generated videos
Hugging Face Datasets2026 · dataset
Pruna Skills doc examplesPruna Skills — doc examples Generated media and sidecars for PrunaAI/pruna-skills documentation. Each file under examples/ matches a skill demo in docs/EXAMPLES.md. PNG/MP3/MP4 outputs have a sibling .meta.json with the exact prompt, model, and inputs. Layout examples/ p-image-advanced.png p-image-advanced.meta.json quickstart-knight-still.png quickstart-knight-clip.mp4 chain-monarch-clip.mp4 … Re
Hugging Face Datasets2026 · dataset · gated
AlayaLab/WildWorldWildWorld: A Large-Scale Dataset for Dynamic World Modelingwith Actions and Explicit State toward Generative ARPG This repo contains the dataset proposed in WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG Zhen Li, Zian Meng, Shuwei Shi, Wenshuo Peng, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang Alaya Studio, Shanda AI Research To
Hugging Face Datasets2026 · dataset
SignLink ASL Video Dictionary MatrixSignLink ASL Video Dictionary Matrix This dataset contains a processed matrix of American Sign Language (ASL) video clips used as the core dictionary lookup engine for SignLink, an offline voice-to-sign inference pipeline. 🤝 Attribution & Data Source The video assets in this dataset are sourced directly from the Sign-Language-Mocap-Archive created by StudioGalt. Original Creator: StudioGalt Source
Hugging Face Datasets2026 · Image
WBenchWBench Dataset A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation TL;DR — WBench evaluates 20 video world models across 5 dimensions and 22 metrics. Overview WBench is a comprehensive multi-turn benchmark for interactive video world model evaluation. It contains 289 multi-turn interaction cases with 1,058 interaction turns for evaluating models across 22 metrics and
Hugging Face Datasets2026 · Text
PhyGroundPhyGround: Benchmarking Physical Reasoning in Generative World Models Project page · Paper · Evaluation code · PhyJudge-9B PhyGround is a criteria-grounded benchmark for diagnosing physical failures in generated video. It contains 250 prompts covering 13 observable physical laws across solid-body mechanics, fluid dynamics, and optics. Each prompt is paired with a first-frame image, 10 released gen
Hugging Face Datasets2026 · dataset
PhysInOnePhysInOne: Visual Physics Learning and Reasoning in One Suite vLAR Group | The Hong Kong Polytechnic University | Syai Singapore | Meta CVPR 2026 🧭 Navigation 📌 Summary 🚀 Release Timetable 📦 Repositories & Downloads 📊 Data Splits 🧱 3D Assets 🛠️ Data Processing 🏆 Leaderboard Evaluation Data 🎞️ Rendered Data 1. Download Scripts 2. Install Dependencies… See the ful
Hugging Face Datasets2026 · dataset · gated
BONES-SEED: Skeletal Everyday Embodiment DatasetBONES-SEED: Skeletal Everyday Embodiment Dataset BONES-SEED is an open dataset of 142,220 annotated human motion animations for humanoid robotics. It provides motion capture data in SOMA and Unitree G1 formats, with natural language descriptions, temporal segmentation, and detailed skeletal metadata. Project website: bones.studio/datasets/seed Interactive viewer: seed-viewer.bones.studio Associate
Hugging Face Datasets2025 · Image
kling v2.1 master Human PreferencesRapidata Video Generation Kling v2.1 Master Human Preference In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Kling v2.1 Master video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our websit
Hugging Face Datasets2025 · Image
KlingTeam/VIVID-10MVIVID-10M [project page] | [Paper] | [arXiv] VIVID-10M is the first large-scale hybrid image-video local editing dataset aimed at reducing data construction and model training costs, comprising 9.7M samples that encompass a wide range of video editing tasks. Data Index The data index is located at four .csv files: vivid-image-change.csv vivid-image-remove.csv vivid-video-change.csv vivid-video-rem
Hugging Face Datasets2025 · Table · CSV
Lixsp11/SekaiSekai: A Video Dataset towards World Exploration This repo contains the dataset proposed in Sekai: A Video Dataset towards World Exploration Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin, Ming Li, Shitian Zhao, Zhaopan Xu, Xinyue Li, Yukang Feng, Jianwen Sun, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Zhixiang Wang, Yuwei Wu, Tong He, Jiangmiao Pang, Yu Qiao, Yunde Jia, Kaipeng Zhang Shanghai AI Labor
Hugging Face Datasets2025 · Image · gated
yuezih/Movie101Movie101 [!NOTE] Please carefully read the Movie101 license before using the data.Current dataset version: Movie101v2 Audio Description (AD) describes movie content in real time to help visually impaired individuals enjoy movies, where a narration speech briefly summarizes the ongoing plots during pauses in character dialogue, help its audience keep up with the movie. The AD creation involves exte
Hugging Face Datasets2025 · Text
wisa-80kWISA-80K Dataset Description WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation Jing Wang*, Ao Ma*, Ke Cao*, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng‡, Yuhui Yin, Xiaodan Liang‡(*Equal Contribution, ‡Corresponding Authors) BibTeX @article{wang2025wisa, title={WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generat
Hugging Face Datasets2025 · Table · CSV
TencentARC/VPDataVideoPainter This repository contains the implementation of the paper "VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control" Keywords: Video Inpainting, Video Editing, Video Generation Yuxuan Bian12, Zhaoyang Zhang1‡, Xuan Ju2, Mingdeng Cao3, Liangbin Xie4, Ying Shan1, Qiang Xu2✉ 1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong 3The University of Tokyo
Hugging Face Datasets2025 · Table · CSV · gated
HOIGen/HOIGen-1MSummary This is the dataset proposed in our paper [CVPR 2025] HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation. HOIGen-1M contains over one million high-quality video clips for HOI video generation with multiple types of HOI videos, diverse scenarios (15, 000+ objects and 7, 000+ interaction types), and expressive captions. HOIGen-1M exhibits three main features: Larg
Hugging Face Datasets2025 · Table · CSV · gated
DropletX/DropletVideo-10M🔍 Dataset Note: DropletVideo-1M is the premium subset of DropletVideo-10M, filtered with aesthetic score > 4.51 and image quality score > 7.51. ✈️ Introduction The challenge of spatiotemporal consistency has long existed in the field of video generation. We have released the open-source dataset DropletVideo-10M —the world's largest video generation dataset with spatiotemporal… See the full descrip
Hugging Face Datasets2024 · Image
Sakugabooru 2025Sakugabooru2025: Curated Animation Clips from Enthusiasts Sakugabooru.com is a booru-style imageboard dedicated to collecting and sharing noteworthy animation clips, emphasizing Japanese anime but open to creators worldwide. Over the years, it has amassed more than 240,000 animation clips, alongside informative blog posts for anime fans everywhere. With the growing interest in generative video mod
Hugging Face Datasets2024 · dataset
OpenVid-1MSummary This is the dataset proposed in our paper [ICLR 2025] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation. OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets. All
Hugging Face Datasets2024 · Table · CSV
Panda-70MPanda 70M dataset by Snap Inc 70M video-caption pairs Code for downloading: https://github.com/snap-research/Panda-70M/dataset_dataloading
Hugging Face Datasets2023 · Image
Pexels-400kPexels 400k Dataset of 400,476 videos, their thumbnails, viewcounts, explicit classification, and caption. Note: The Pexels-320k dataset in the repo is this dataset, with videos <10s removed.