{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:image-text-to-text",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Image
OneJev-DataThe training data of OneJev: 94,707 typed questions about screens, photos, videos and text, each with its answer. This is 95.5% of the rows OneJev was trained on; rows whose sources do not allow redistribution are left out. Use from datasets import load_dataset ds = load_dataset("OmniJev/OneJev-Data", split="train", streaming=True) row = next(iter(ds)) Each row has a state with <image:N> and <vide
Hugging Face Datasets2026 · Image
ImajevBench v2.0-lite (preview)ImajevBench v2.0-lite (preview) A benchmark for typed decisions made from a photo, a written rule, or both. A system receives the evidence (0–2 images and a state: an order, a rule, a form, a claim) and one typed question, and must return exactly one of: yes/no, a listed choice, an integer level, or Unknown when the evidence does not determine the answer. It checks four things: whether the decisio
Hugging Face Datasets2026 · Image
DAREBench: Deployment-Aware and Reliable Evaluation of Models as AgentsDAREBench DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents) is a workload- and deployment-aware benchmark for evaluating models as agents. Built on a shared OpenClaw execution environment, it organizes 233 tasks selected and adapted from 22 source benchmarks into a 2×3 workload matrix defined by input modality and execution form, and evaluates them under a unified contract-b
Hugging Face Datasets2026 · Image
AgroOmniAgroOmni AgroOmni is a large-scale multi-view agricultural dataset introduced in "AgroOmni: A Large-Scale Multi-view Agricultural Dataset for Cross-Scale Multimodal Reasoning" (arXiv:2603.14342). It spans ground, UAV, and satellite imagery to address the ground-level bias of existing agricultural multimodal models, with 288K visual question answering pairs covering 56 specialized task categories a
Hugging Face Datasets2026 · Image
Omni-Edu (Core V6)Omni-Edu — Core V6 SFT mixture 69,999 supervised instruction examples (~158M characters) covering K-12 subject competence, curriculum grounding, diagnostic reasoning, pedagogical action and general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image referenced by the JSONL ships in this repository under images/. This is the system-prompted assembly of the v6 core mixture: every ro
Hugging Face Datasets2026 · Text
The best of SWEDataset Description This dataset is a filtered and deduplicated version of a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 291.74kb, a total uncompressed size of 14.93GB
Hugging Face Datasets2026 · Text
The best of SWEDataset Description This dataset is a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 169.08kb, a total uncompressed size of 19.80GB, and a total of 1227
Hugging Face Datasets2026 · Image
rsoohyun/SpatialBlock-15kSpatialBlock-15k This dataset accompanies the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem. It contains 15,000 synthetic block-stacking problems for training large vision-language models (LVLMs) to improve spatial reasoning. The dataset includes three types of multiple-choice questions: Q1: 3D-to-2D projection Q2: viewpoint transformation Q3: str
Hugging Face Datasets2026 · Image · gated
PopResumePopResume Population-Representative Resume Dataset for Causal Fairness Evaluation of LLM/VLM Resume Screeners 🌐 Project page · 📄 arXiv · 💻 Code · EMNLP 2026 Main (Oral) PopResume is a synthetic, population-representative resume dataset that preserves the natural statistical relationships among demographic attributes, demographic proxies, and job-relevant qualifications found in U.S. population sta
Hugging Face Datasets2026 · Image
OPPOer/IC-VCO-DatasetIC-VCO-Dataset This dataset package contains the two IC-VCO training subsets: sft: supervised fine-tuning examples. preference: visual contrastive preference examples. The two subsets intentionally use different schemas, so they are represented as separate Hugging Face dataset configurations instead of separate splits under a single configuration. Each configuration has a train split
Hugging Face Datasets2026 · Image
Kairos Multimodal ReasoningA dataset for training models in multimodal reasoning tasks Usage from datasets import load_dataset ds = load_dataset("Aquiles-ai/Kairos-Multimodal-Reasoning") print(ds.features) print(ds["train"]["source"]) Preview of dataset examples We've built a playground so you can see some of the examples included in the dataset. Link: https://kairos-example.vercel.app/ Dataset used in the blog post: Kairos
Hugging Face Datasets2026 · Image
lenamerkli/distilled-webDataset Card for lenamerkli/distilled-web This dataset consists of web-scraped data using a custom crawler purpose-built for each website. Dataset Details Dataset Sources Repository: https://github.com/lenamerkli/distilled-web Uses This dataset is useful for training large language models. The train split provides instruction-following and chat data for supervised fine-tuning (SFT) and instruction
Hugging Face Datasets2026 · dataset
DCVLM-Pool (large)DCVLM-Pool (large) The raw candidate pool at the large scale of our DataComp-VLM benchmark: 1,949,321,868 samples / 166.7 TB across 166 source datasets, as WebDataset tar shards — ≈4× the medium pool. 🚚 Upload in progress This repo is being populated incrementally and is not yet complete — shards are still being uploaded. Sources already present are final and safe to use; sources with fewer shards
Hugging Face Datasets2026 · dataset
baochenfu/OCT-Bench🤗 OCT-Bench Paper | GitHub We introduce OCT-Bench, a comprehensive benchmark for evaluating Multimodal Large Language Models (MLLMs) on optical coherence tomography (OCT) image understanding. OCT-Bench comprises 10,076 expert-verified multiple-choice questions from 4,137 OCT images across seven public datasets and evaluates 3 capability dimensions, 9 capability groups, and 20 fine-grained tasks co
Hugging Face Datasets2026 · Image
Tianhai266666/PowerCoT-1PowerCoT-1 PowerCoT-1 is an instruction-tuning dataset for power-system dispatch and control. It contains multimodal image-text reasoning samples and general text reasoning samples, each annotated with a chain-of-thought (CoT). The dataset is built through a multi-stage data-engineering pipeline with rule-based cleaning and manual spot-checks, totaling roughly 1.8 million samples. Dataset Structur
Hugging Face Datasets2026 · Text
AnomalyThinkAnomalyThink: reasoning traces for explainable industrial anomaly detection AnomalyThink is a collection of structured reasoning traces for industrial anomaly detection (IAD), distilled from Gemini 2.5-Flash on Real-IAD images. Each example is a single-image inspection in which the assistant produces a <think> reasoning trace, a defect <location> and <type> (for anomalies), and a binary <answer> (
Hugging Face Datasets2026 · Image · gated
Dress-EDDress-ED Instruction-Guided Editing for Virtual Try-On and Try-Off Overview Dress-ED is a large-scale garment-editing dataset for instruction-guided virtual try-on and clothing manipulation research. Given a person image and a garment image, each example asks a model to apply a specific textual edit to the garment while keeping it worn on the person. The dataset provides three complementary instru
Hugging Face Datasets2026 · Text
build-small-hackathon/blood-test-explainer-tracesBlood Test Explainer - agent traces Agent traces from the Blood Test Explainer app (Build Small hackathon). Each row is one publicly-available sample lab report (fake patients, no PHI) run through the full agent pipeline: a small vision model reads the document and extracts the markers, then a curated medical knowledge base turns the values into a grounded, per-marker explanation plus cross-marker
Hugging Face Datasets2026 · Image
AiTW Processed Full with App LabelsAiTW Processed Full with App Labels This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset. Why This Exists AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and act
Hugging Face Datasets2026 · Text
Manga109-s Text Line AnnotationsManga109-s Text Line Annotations High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP). Notice: This d
Hugging Face Datasets2026 · Image
PCF-BenchPCF-Bench A photonic-crystal-fiber (PCF) inverse-design benchmark for vision-language models. Each sample bundles geometric parameters, simulated mode-field images, and a four-level expert-style annotation suite supporting tasks across geometry perception, physics understanding, multimodal reasoning, inverse design, and code generation. Note (anonymous review). This dataset card omits identifying
Hugging Face Datasets2026 · Text
NoRANoRA Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning Paper | Code and model interface NoRA evaluates the actions a model proposes, the facts it observes, and the reasons connecting them in first-person scenes. Dataset Split Annotations Clips Facts Reasons Actions train LLMSilver: model-annotated 1,230 8,731 7,101 4,423 test HumanGold: human-reviewed 190 1,324 1
Hugging Face Datasets2026 · Image · gated
Nalandadata/nalanda-image-qaNalanda Image QA 22,679 multimodal STEM science question-answer pairs with diagrams and chain-of-thought answers. Used to train Nalandadata/nalanda-image-vl — fine-tuning LLaMA-3.2-Vision-11B raised accuracy from 37.7% to 50.0% (+12.3 points) on the held-out evaluation set. 🏆 Live Leaderboard: Nalanda Image VL Leaderboard — see how frontier models rank on this benchmark. 📦 Public sample (no login
Hugging Face Datasets2026 · Text
Nemotron Image Training v3Nemotron Image Training v3 Versions Date Commit Changes 2026-04-28 HEAD Initial commit. Dataset Description Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card de
Hugging Face Datasets2026 · dataset
internlm/WildClawBenchWildClawBench Hard, practical, end-to-end evaluation for AI agents — in the wild. WildClawBench is an agent benchmark that tests what actually matters: can an AI agent do real work, end-to-end, without hand-holding? We drop agents into a live OpenClaw environment — the same open-source personal AI assistant that real users rely on daily — and throw 60 original tasks at them: clipping goal highligh