{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:image-classification",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · dataset
Herculaneum legibility V5 — training data and recipeHerculaneum legibility V5 — training data and recipe (R2B, EXPERIMENTAL) This is the cohort, splits, source bundles, raw crops and full code behind the V5 scorer: https://huggingface.co/LimeGS/herculaneum-legibility-proxy-v5. V5 is a new recipe that improves on the published proxy_v4 (https://huggingface.co/LimeGS/herculaneum-legibility-proxy). It retrains a ResNet-18 on a corrected historical coh
Hugging Face Datasets2026 · Image
DOCO ImageNet-L and OOD-L with LAION-C DistortionsDOCO ImageNet-L and OOD-L with LAION-C Distortions This repository contains the materialized ImageNet-L and OOD-L image sets used for the LAION-C experiments in Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensation (DOCO) (CVF PDF, arXiv:2604.21772, code). The artifacts were generated by applying a modified version of the LAION-C corruption pipeline to ImageNet validation
Hugging Face Datasets2026 · Image
荆楚文化文物语义分析样本荆楚文化文物语义分析样本 这是用于审核字段设计和语义抽取质量的样本版本,共 116 条记录、35 个核心字段。 数据集另含 enrichment_trial_5 配置:从主表选取 5 条文物进行检索、图像观察与语义补充,共 40 个字段。主表内容未被覆盖。 湖北省博物馆:83 条 荆州博物馆:33 条 图片位于第 2 字段 image_url 两个古籍书影汇总页已拆分为 18 条单书记录 删除了当前来源完全无法填充的 creation_place 和 collection_number 删除派生检索字段 keywords;删除与 archaeological_site 高度重复的 provenance 原文与语义归纳分离;缺少来源的信息保持为空 所有记录目前均为 待人工复核,尚不是最终 604 条全量版本 数据文件 data/artifacts.csv:Dataset Viewer 使用的
Hugging Face Datasets2026 · dataset · gated
AIGC Dataset — NovelAI GenerationsAIGC Dataset — NovelAI Generations ⚠️ Contains NSFW content. Images and prompts have not been filtered. Many are sexually explicit. 1,562,685 AI-generated images (≈332 GB) made with NovelAI's image models (V3 → V4 → V4.5 → V5) between 2023-11 and 2026-09 through a private Discord bot (Kohaku-NAI). Every image has its full generation metadata (model, sampler, steps, CFG, seed, prompt, …) and an exa
Hugging Face Datasets2026 · Magnetic resonance imaging · gated
OAIDataset Card for OAI Knee MRI (Preprocessed 3D NumPy Arrays) Dataset Details Dataset Description This dataset contains 92,313 verified, preprocessed Knee MRI volumes from the Osteoarthritis Initiative (OAI). The original DICOM/medical imaging files have been standardized, cleaned, and converted into highly efficient 3D .npy (NumPy) arrays to accelerate machine learning and computer vision workflow
Hugging Face Datasets2026 · Image
PolyMFOPolyMFO PolyMFO is a high-resolution image dataset for foreign-object anomaly classification and binary segmentation. It contains normal samples and samples from nine foreign-object classes. Every image has a corresponding binary mask. This Hugging Face repository contains the dataset only.Source code for data processing, benchmark preparation, evaluation, and related research utilities is maintai
Hugging Face Datasets2026 · Image · gated
Car Damage ImagesCar Damage Images A raw image collection for vehicle damage assessment. Unlabeled: these images have no annotations yet and are intended as source material for labeling or pre-training. Structure images/ 001/ images_001.jpg images_002.jpg ... 002/ ... ... 683 folders thumbnails/ 001/ thumbnail_001.jpg thumbnail_002.jpg ... ... 683 folders Branch Files Folders Size… See the full description on the
Hugging Face Datasets2026 · Computed tomography
vectorsense/organscan-dataOrganScan training corpus — recipe and provenance This repository publishes the recipe, not the pixels. Medical imaging corpora carry licences that differ source by source, and several forbid redistribution outright. More importantly, frames rendered from DICOM can carry burned-in patient identifiers even after tag-level de-identification. Rather than ship images that are easy to misuse, this repo
Hugging Face Datasets2026 · Image
PESTPEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation 🎉 ACM Multimedia (ACM MM) Main Track Paper Accepted! Figure 1. Overview of the PEST framework. 🔎 TL;DR PEST introduces a novel framework for generalizable hateful meme moderation that leverages lightweight task-specific multimodal agents to steer powerful black-box vision-language models
Hugging Face Datasets2026 · Image
GTSRB for PranavX AI Reliability ExperimentsGTSRB for PranavX AI Reliability Experiments This repository repackages the German Traffic Sign Recognition Benchmark (GTSRB) for a traffic-sign classification reliability experiment. It preserves the original PPM image bytes and GTSRB class IDs in Parquet. No image is resized, cropped further, enhanced, or synthetically corrupted here. This is a project-specific derivative split, not a new offici
Hugging Face Datasets2026 · Image
Synthetic Text Images (English)Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise.
Hugging Face Datasets2026 · dataset
CatsVsDogImgClsDatasetThis is clone of popular Cats Vs. Dogs image classification dataset for CNN-based visual classification models. If you are using raw images (stored in /data), this script can be used for creating train-test split on the raw images dataset: def fetching_selected_data(cats_images_subpath: str, dogs_images_subpath: str, test_split: int, max_len: int = 1000000000): total_len = min(len(os.listdir(cats_
Hugging Face Datasets2026 · Image
Britannica Illustrated PagesBritannica Illustrated Pages 115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition (1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes (838 Internet Archive items). A second config carries the classifier score, OCR word count and provenance for every one of the 975,345 pages. Two things the scan showed: 82% of the illust
Hugging Face Datasets2026 · Image
Biomedica 2025 Eval SetBiomedica 2025 Eval Set Unified test snapshot of the BioMedica 2025 vision–language evaluation suites used in AMInZeroShotOpenEvalAllTasks. Every row is a single image with closed-ended options, the gold answer, and provenance fields that point back to the original dataset. Images are stored as original JPEG/PNG bytes (or JPEG-encoded arrays) inside parquet so the Hugging Face dataset viewer is en
Hugging Face Datasets2026 · Image
tomato_leavesTomato Leaves Dataset Overview This dataset contains images of tomato leaves categorized into different classes based on the type of disease or health condition. The dataset is divided into training, validation, and test sets, with a ratio of 8:1:1. The classes include various diseases as well as healthy leaves. The dataset includes both augmented and non-augmented images. Dataset Structure The da
Hugging Face Datasets2026 · dataset
danbooru2025Danbooru2026: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset [WIP] Dataset Description Danbooru2026 is a large-scale anime illustration dataset containing over 10 million community-annotated images. It is intended for research and development in anime-style image generation, image classification, multimodal learning, and related tasks. Danbooru is a long-running image board known
Hugging Face Datasets2026 · Image
DatarrX/pyu-handwritten-consonant-datasetMyanmar’s Ancient Heritage: Pyu Handwritten Consonant Dataset An open-access, systematically curated handwritten dataset of the 33 ancient Pyu consonants. This project serves as a foundational baseline benchmark to support digital humanities, paleographical preservation, and advanced computer vision tasks such as Optical Character Recognition (OCR). The dataset is modeled directly after canonical
Hugging Face Datasets2026 · Image
SPARK-2022SPARK 2022 — Stream 1 (Spacecraft Detection) Stream 1 of the SPARK 2022 dataset (SPAcecraft Recognition leveraging Knowledge of the space environment): space-borne imagery of 10 spacecraft plus a debris class, for object detection and classification. Each image contains exactly one target annotated with a single bounding box and class label. Dataset summary Images 110,000 JPEG, 1024 × 1024, RGB An
Hugging Face Datasets2026 · Image
MetaPKLotMetaPKLot A Large-Scale Benchmark for Vision-Based Parking Lot Management 2,265,974 labeled samples · 1,366,185 new annotations · 3 research challenges · COCO-style annotations MetaPKLot is a large-scale, harmonized dataset designed for research on vision-based parking lot management. It extends and standardizes three existing parking datasets: PKLot CNRPark-EXT PLds MetaPKLot introduces new annot
Hugging Face Datasets2026 · Image
Vasanthaleela/engineering-drawings-as1100Engineering Drawings AS1100 Compliance Dataset Dataset Description This dataset contains engineering drawings with various AS1100 (Australian Standard for Technical Drawing) compliance issues for training AI models to identify missing elements and non-compliance issues in technical drawings. Dataset Summary The Engineering Drawings AS1100 Compliance Dataset is designed to train and evaluate vision
Hugging Face Datasets2026 · Image
SVG Generation Benchmark (Static)Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,355,161 human responses, collected with the Rapidata Python SDK, comparing how well 30 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
Hugging Face Datasets2026 · Image
OculoOculo Oculo is a publicly available multi-label ocular B-scan ultrasound dataset of 1,630 images from 1,242 patients, annotated for five ophthalmic abnormalities. It accompanies the MICCAI 2026 paper Oculo: A Multilabel Dataset for Identification of Ocular Abnormalities from Ultrasound Images. Benchmark code: https://github.com/Sri-Kanchi-Kamakoti-Medical-Trust/Oculo Paper: https://papers.miccai.o
Hugging Face Datasets2026 · Image
SVG Generation Benchmark (Static)Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
Hugging Face Datasets2026 · dataset
Children Gait VideoDecoding Children's Gait Behavior ECCV 2026 Yifan Shen1,2,*, Boyi Li1,*, Meihuan Huang2,3,4,*, Yuanzhe Liu1,*, Xu Cao1,2,*,§, Jinyang Jin1, Zhengyuan Li1, Anglin Liu5, Junho Kim1, Jingyuan Zhu2, Fangzhou Lan2, Jianguo Cao2,3, Jintai Chen5, Ismini Lourentzou1, James M. Rehg1,† 1 University of Illinois Urbana-Champaign 2 PediaMed AI 3 Shenzhen Children's Hospital 4 Hong Kong Polytechni
Hugging Face Datasets2026 · Computed tomography
CancerVerse🩻 CancerVerse The first longitudinal, multimodal CT dataset spanning 13 malignant tumor types CancerVerse pairs whole-body abdominal/pelvic CT volumes with the radiologists' own free-text reports, follows patients across time, and is being released in stages — culminating in expert voxel-level tumor masks and a deep clinical & longitudinal annotation layer. It is built to power the next generation