{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:object-detection",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · dataset
DragonCheat-AI/DragonData-4Dragon 4 Dataset Dataset for object detection in Standoff 2. Contains labeled screenshots with two classes: ct (Counter-Terrorist) and tr (Terrorist). Contents Format: YOLO (images + .txt labels) Classes: ct, tr Split: train / valid / test License MIT
Hugging Face Datasets2026 · Image
PolyTopoBenchPolyTopoBench PolyTopoBench is a benchmark for complex vector polygon generation from remote-sensing imagery. It evaluates whether a model can produce complete polygon topology, including interior rings (holes), rather than only exterior boundaries. Paper: PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery, NeurIPS 2026 (Evaluations and Datasets Track) Cod
Hugging Face Datasets2026 · dataset
UrbanAnonymizer DatasetUrbanAnonymizer Dataset Mehmet Kerem Turkcan Columbia University The UrbanAnonymizer Dataset contains 134,444 images with 394,800 face and license plate boxes in COCO format, compiled from Open Images, WIDER FACE, DARK FACE, CRPD and BirdsEye-RU. Each image lists the classes whose annotation is complete (verified_classes), so detectors can be trained and evaluated on partially labeled sources with
Hugging Face Datasets2026 · Image
KeenForgeAI/NEU-DET-correctedNEU-DET-corrected Hot-rolled steel strip surface defect detection — cleaned version of the NEU Surface Defect Database. 热轧带钢表面缺陷检测 —— NEU 表面缺陷数据库的清理修正版。 English · 中文 English What is this? A cleaned, complete and consistently packaged version of the NEU Surface Defect Database (NEU-DET) from Northeastern University (Song Kechen & Yan Yunhui). The upstream release ships as a flat IMAGES/ + ANNOTATIO
Hugging Face Datasets2026 · Image
Mira-Scene DatasetMira-Scene Dataset This repository contains the public data used by the Mira-Scene single-image 3D scene reconstruction project: the BlendSwap evaluation benchmark and the training releases derived from Objaverse Outpaint and 3D-FRONT. The benchmark and training data are published in the same Hugging Face dataset repository. The training data is distributed as verified tar.gz shards. The archives
Hugging Face Datasets2026 · Image
miti360Miti360: An integrated dataset combining remote sensing, ground measurements and weather data for improved reforestation monitoring Introduction In the era of artificial intelligence, machine learning combined with remote sensing and ground measurements offers unprecedented opportunities to enhance forest monitoring through faster, more accurate biomass estimation and individual tree analysis. Des
Hugging Face Datasets2026 · Image
RA-4MRA-4M — a verified free-text relation corpus 472,344 images · 4,282,531 relations · 10,102 free-text predicates · 497 object categories · 9.03 relations per image in the training split, plus a 24,964-image validation split (226,203 relations). Machine-annotated by an open-weight vision–language model, then filtered by deterministic geometric gates that reject 11.3% of raw candidates. Every object
Hugging Face Datasets2026 · Image
AI-TOD-v2AI-TOD-v2 AI-TOD-v2, the tiny-object detection benchmark in aerial images, packed once with the official v2 annotations kept whole, so it loads in one line and no data path has to be configured: from datasets import load_dataset ds = load_dataset("shijli/aitod-v2") # 11214 train / 2804 validation / 14018 test AI-TOD cuts 28036 images of 800 x 800 pixels from xView, DOTA-v1.5, VisDrone2018-Det, Air
Hugging Face Datasets2026 · Image · gated
Dinoman1221/sonarvision-multisource-v6🌊 KADAL Multi-Source Side-Scan Sonar Dataset (v6) SIH 2026 | SIH26057 (Ministry of Earth Sciences / NIOT) — AI-Powered Automated Underwater Marine Debris and Anomaly Detection System using Side-Scan Sonar Imagery. Official Project Source Code & Documentation:🔗 https://github.com/Dinoman67/sonarvision(Refer to the GitHub repository for preprocessing scripts, augmentations, model weights, edge dashb
Hugging Face Datasets2026 · Image
PDFA OCR Dataset - KREATIVE TIME BOXPDFA OCR Dataset Curated and Published by KREATIVE TIME BOX This dataset contains document page images along with their corresponding OCR layout bounding box annotations derived from PDFA document extraction pipelines. Dataset Overview Organization / Creator: KREATIVE TIME BOX Images: 27,499 PNG files (~8.8 GB) JSON Annotations: 6,989 JSON files (~52 MB) Image Format: PNG (RGB document page render
Hugging Face Datasets2026 · dataset · gated
SoccerTrack v2SoccerTrack v2 Ten university-level soccer matches (934 minutes) recorded by fixed panoramic camera systems whose field of view spans the entire pitch, released together with frame-level game state annotations and player-linked ball action events on the same footage. Contents Folder Contents videos/ 20 panoramic half-match videos (two 45-minute periods per match) gsr/ Game state annotations: one J
Hugging Face Datasets2026 · Image
Indonesian KTP Dataset 24K (Flat & 3D Perspective Augmented)Indonesian KTP Dataset 24K (Flat & Augmented - Commercially Safe) Welcome to the Indonesian KTP (Kartu Tanda Penduduk) Dataset. This is a highly robust, high-fidelity, and commercially safe synthetic dataset designed to advance SOTA (State-of-the-Art) research in Document Information Extraction (DIE), Key Information Extraction (KIE), and Optical Character Recognition (OCR) specifically for Indone
Hugging Face Datasets2026 · dataset
LocateAnything-DataLocateAnything-Data 中文 · Paper · Model · Code Overview LocateAnything-Data is the public training-data release for LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding. LocateAnything formulates detection and visual grounding as a unified vision-language task. Given an image and a category, phrase, text string, or action-oriented instruction, the model predict
Hugging Face Datasets2026 · Image
FZI-AURAFZI-AURA A multimodal autonomous-driving dataset featuring the largest LiDAR sensor suite of any public autonomous-driving dataset. Download | Python SDK | Data format | License 2,473 scenes | 13.70 hours | 8 cameras | up to 12 LiDARs | 4.11 million 3D-box annotations | 30.01 billion segmented LiDAR points Public preview release: 1,081 out of the 2,473 FZI-AURA scenes are currently available. The
Hugging Face Datasets2026 · Text · gated
PathBindPathBind A diagnostic benchmark for evaluating pathology vision-language models. PathBind bundles 2,600 samples across three components — each filtered by an automated pipeline and finalized under expert pathologist review — to jointly probe visual dependence, cross-domain replication, and entity-level visual-semantic binding. Config # samples Format PathBind-VQA 1,500 sample-ID manifest (TSV) Pat
Hugging Face Datasets2026 · Image
DatarrX/pyu-handwritten-consonant-datasetMyanmar’s Ancient Heritage: Pyu Handwritten Consonant Dataset An open-access, systematically curated handwritten dataset of the 33 ancient Pyu consonants. This project serves as a foundational baseline benchmark to support digital humanities, paleographical preservation, and advanced computer vision tasks such as Optical Character Recognition (OCR). The dataset is modeled directly after canonical
Hugging Face Datasets2026 · Image
SPARK-2022SPARK 2022 — Stream 1 (Spacecraft Detection) Stream 1 of the SPARK 2022 dataset (SPAcecraft Recognition leveraging Knowledge of the space environment): space-borne imagery of 10 spacecraft plus a debris class, for object detection and classification. Each image contains exactly one target annotated with a single bounding box and class label. Dataset summary Images 110,000 JPEG, 1024 × 1024, RGB An
Hugging Face Datasets2026 · Image
biglam/artinsight-painting-deteriorationArtInsight — Easel Painting Deterioration Detection 20 high-resolution full-frame photographs of easel paintings, annotated by expert restorers with the areas of damage they see. 2,909 annotations in total. Conservation assessment is normally done by eye, by specialists, one painting at a time. This dataset is an attempt to make it learnable. Two damage types, two disjoint sets of paintings Subset
Hugging Face Datasets2026 · Image
MetaPKLotMetaPKLot A Large-Scale Benchmark for Vision-Based Parking Lot Management 2,265,974 labeled samples · 1,366,185 new annotations · 3 research challenges · COCO-style annotations MetaPKLot is a large-scale, harmonized dataset designed for research on vision-based parking lot management. It extends and standardizes three existing parking datasets: PKLot CNRPark-EXT PLds MetaPKLot introduces new annot
Hugging Face Datasets2026 · Image
HEIR: Learning Human-Entity Interactions with Functional RolesHEIR: Learning Human-Entity Interactions with Functional Roles LOCAL DRAFT — NOT CLEARED FOR PUBLICATION. Image redistribution evidence and source attribution are incomplete. The annotation license has not been finalized. See the local release privacy audit before uploading this directory. Code, authors and citation · Download guide HEIR represents each person–action event as a complete set of par
Hugging Face Datasets2026 · Image
llama-farm/military-labeled-yoloMilitary-Labeled YOLO Dataset (DVIDS sourced) Multi-class military object detection in YOLO format. Source images pulled from the Defense Visual Information Distribution Service (DVIDS) public domain library; labeled via in-house Gemini-VLM-assisted pipeline with human-in-the-loop correction. Classes (12) ID Name 0 soldier 1 tank 2 apc 3 artillery 4 mlrs 5 military_truck 6 helicopter 7 aircraft 8
Hugging Face Datasets2026 · Text
Manga109-s Text Line AnnotationsManga109-s Text Line Annotations High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP). Notice: This d
Hugging Face Datasets2026 · Image
ROD-Dataset — Real-Time Obstacle Detection (YOLO)ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision 24,326-image, 25-class YOLO dataset for obstacle detection This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living w
Hugging Face Datasets2026 · Table · Parquet
PubMed-OphthaPubMed-Ophtha PubMed-Ophtha is a hierarchical ophthalmic vision-language dataset built from open-access articles in PubMed Central. Figures are extracted directly from the article PDFs at full resolution and decomposed into panels, panel identifiers, and individual images; figure captions are split into panel-level subcaptions. The result is a panel-centric corpus of 102,023 panels from 35,544 fig
Hugging Face Datasets2026 · Image
ai4data/data-snapshotDataset card for data-snapshot This dataset was introduced in the paper Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents. The source code for the benchmark and dataset extraction is available on GitHub: worldbank/ai4data. Dataset summary The data-snapshot dataset is an annotated corpus designed for the evaluation and development of models f