Hugging Face Datasets2026 · Table · CSV
LIFT-VistaLIFT-Vista LIFT-Vista is a dataset with large camera viewpoint changes and joint camera-layout annotations. Overview camera.csv camlayout.csv Clips 120,898 58,272 (a subset of the camera clips) Annotations camera trajectory, caption camera trajectory, caption, last-frame layout, per-frame box tracks Every clip has 81 frames at 16 fps (about 5 s) at the native resolution of its source video (98% ar
Hugging Face Datasets2026 · Image
VidScribeVidScribe VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12). Release status. This repository hosts the public hal
Hugging Face Datasets2026 · Image
Relative Camera Movement Benchmark (Image-to-Video)Rapidata Relative Camera Movement Benchmark Built by Rapidata. This dataset contains 900,318 human responses, collected with the Rapidata Python SDK, comparing how well 15 image-to-video models and world models move the camera relative to what is in the scene. Each row is a head-to-head comparison between two models' clips generated from the same still and the same instruction, judged by human ann
Hugging Face Datasets2026 · dataset
Object PermanenceObject Permanence The training corpus of WROP (World Reasoning with Object Permanence): 1.5M Blender-rendered video-continuation samples across 150 hand-designed cognitive tasks, one tar per task. Abstract Object permanence is the hallmark of human cognitive priors. Recent studies show that video models… See the full description on the dataset page: https://huggingface.co/datasets/Hokin/object-per
Hugging Face Datasets2026 · dataset · gated
Omni-R2V DatasetOmni-R2V Dataset Large-scale training data for omni reference-to-video generation 🔊 Audio-enabled target videos. Audio tracks are retained where available, enabling audio-bearing samples to support reference-to-audiovisual (R2AV) training alongside R2V. Overview · Task coverage · Quick start · Data format · Citation From individual reference factors to multi-content and cross-aspect combinations.
Hugging Face Datasets2026 · Image
Camera Movement Benchmark (Image-to-Video)Rapidata Camera Movement Benchmark Built by Rapidata. This dataset contains 324,044 human responses, collected with the Rapidata Python SDK, comparing how well 15 image-to-video models and world models execute a described camera movement from a single still image. Each row is a head-to-head comparison between two models' clips generated from the same still and the same instruction, judged by human
Hugging Face Datasets2026 · dataset
Magpie Dataset LiteMagpie Dataset Lite Paper: Magpie: Real-Time World Renderer for Interactive GamesProject Page: https://zhanxy.xyz/Magpie-website Magpie Dataset Lite is a publicly released subset of the Magpie interactive game rendering dataset (arXiv:2608.27168). Magpie is a real-time generative world-rendering system that separates gameplay execution in a game engine from visual synthesis in a render server. Thi
Hugging Face Datasets2026 · Image
VideoArgusBenchVideoArgusBench Sample-specific rubric benchmark for conditioned video generation. VideoArgusBench is the evaluation benchmark for VideoArgus, a framework that scores a generated video against a rubric written for that specific prompt rather than a fixed global metric. This dataset ships the inputs (conditioning assets + prompts) and, for each input, a rubric. It does not contain generated videos
Hugging Face Datasets2026 · dataset · gated
AlayaLab/WildWorldWildWorld: A Large-Scale Dataset for Dynamic World Modelingwith Actions and Explicit State toward Generative ARPG This repo contains the dataset proposed in WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG Zhen Li, Zian Meng, Shuwei Shi, Wenshuo Peng, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang Alaya Studio, Shanda AI Research To
Hugging Face Datasets2026 · Image
WBenchWBench Dataset A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation TL;DR — WBench evaluates 20 video world models across 5 dimensions and 22 metrics. Overview WBench is a comprehensive multi-turn benchmark for interactive video world model evaluation. It contains 289 multi-turn interaction cases with 1,058 interaction turns for evaluating models across 22 metrics and
Hugging Face Datasets2026 · Text
PhyGroundPhyGround: Benchmarking Physical Reasoning in Generative World Models Project page · Paper · Evaluation code · PhyJudge-9B PhyGround is a criteria-grounded benchmark for diagnosing physical failures in generated video. It contains 250 prompts covering 13 observable physical laws across solid-body mechanics, fluid dynamics, and optics. Each prompt is paired with a first-frame image, 10 released gen
Hugging Face Datasets2026 · dataset
PhysInOnePhysInOne: Visual Physics Learning and Reasoning in One Suite vLAR Group | The Hong Kong Polytechnic University | Syai Singapore | Meta CVPR 2026 🧭 Navigation 📌 Summary 🚀 Release Timetable 📦 Repositories & Downloads 📊 Data Splits 🧱 3D Assets 🛠️ Data Processing 🏆 Leaderboard Evaluation Data 🎞️ Rendered Data 1. Download Scripts 2. Install Dependencies… See the ful
Hugging Face Datasets2026 · dataset · gated
Hi-SingersHi-Singers A Comprehensive High-Quality Dataset for Expressive Audio-Driven Singing Head Synthesis Paper / DOI · Supplementary material · Browse files · License · Citation [!IMPORTANT] Research use only. Hi-Singers is a gated dataset released for non-commercial academic research. Access requires acceptance of the Hi-Singers Research License. Redistribution of the raw data and attempts to identify,
Hugging Face Datasets2026 · Table · Parquet
linyq/kiwi_edit_training_dataRefVIE (Kiwi-Edit Training Data) Project Page | Paper | GitHub RefVIE is a large-scale dataset tailored for instruction-reference-following video editing tasks, introduced in the paper "Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance". The dataset was constructed using a scalable data generation pipeline that transforms existing video editing pairs into high-fidelity trai
Hugging Face Datasets2026 · dataset
DAGroup-PKU/RoVid-XRethinking Video Generation Model for the Embodied World If you like our project, please give us a star ⭐ on GitHub for the latest update. Key features 4M robotic video clips(10K+ hours) for large-scale video generation training. 1300+ fine-grained robotic skills, covering diverse actions and task primitives. Multi-modal physical annotations, including RGB, depth, and optical flow. Multi-robot and
Hugging Face Datasets2025 · Image · gated
DNA-Rendering-ProcessedDNA-Rendering-Processed Dataset Project Page | Paper | Code | Model To enable Diffuman4D model training, we meticulously process the DNA-Rendering dataset by recalibrating camera parameters, optimizing image color correction matrices (CCMs), predicting foreground masks, and estimating human skeletons. To promote future research in the field of human-centric 3D/4D generation, we have open-sourced o
Hugging Face Datasets2025 · Table · CSV
Lixsp11/SekaiSekai: A Video Dataset towards World Exploration This repo contains the dataset proposed in Sekai: A Video Dataset towards World Exploration Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin, Ming Li, Shitian Zhao, Zhaopan Xu, Xinyue Li, Yukang Feng, Jianwen Sun, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Zhixiang Wang, Yuwei Wu, Tong He, Jiangmiao Pang, Yu Qiao, Yunde Jia, Kaipeng Zhang Shanghai AI Labor
Hugging Face Datasets2025 · Table · CSV
MuteApo/RealCam-VidRealCam-Vid Dataset News 25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid. 25/03/26: Release our dataset RealCam-Vid v1 for metric-scale camera-controlled video generation, containing ~100K video clips with dedicated short/long captions and metric-scale camera annotations. 25/02/18: Initial commit of the project, we plan to release the full dataset and data process
Hugging Face Datasets2025 · Table · CSV
TencentARC/VPDataVideoPainter This repository contains the implementation of the paper "VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control" Keywords: Video Inpainting, Video Editing, Video Generation Yuxuan Bian12, Zhaoyang Zhang1‡, Xuan Ju2, Mingdeng Cao3, Liangbin Xie4, Ying Shan1, Qiang Xu2✉ 1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong 3The University of Tokyo
Hugging Face Datasets2025 · Table · CSV · gated
DropletX/DropletVideo-10M🔍 Dataset Note: DropletVideo-1M is the premium subset of DropletVideo-10M, filtered with aesthetic score > 4.51 and image quality score > 7.51. ✈️ Introduction The challenge of spatiotemporal consistency has long existed in the field of video generation. We have released the open-source dataset DropletVideo-10M —the world's largest video generation dataset with spatiotemporal… See the full descrip
Hugging Face Datasets2024 · dataset
OpenVid-1MSummary This is the dataset proposed in our paper [ICLR 2025] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation. OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets. All
Hugging Face Datasets2024 · Table · CSV
Panda-70MPanda 70M dataset by Snap Inc 70M video-caption pairs Code for downloading: https://github.com/snap-research/Panda-70M/dataset_dataloading
Hugging Face Datasets2023 · Image
Pexels-400kPexels 400k Dataset of 400,476 videos, their thumbnails, viewcounts, explicit classification, and caption. Note: The Pexels-320k dataset in the repo is this dataset, with videos <10s removed.