{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:zero-shot-classification",
"evidence": null
}
],
"evidence_policy": "broad"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Table · CSV
Finance Text ClassificationFinance Text Classification Dataset A lightweight English-language dataset designed for text classification tasks in the finance domain.It contains 1,000 synthetic but realistic loan application texts with requested amounts up to €50,000. The dataset is suitable for: Binary text classification (approved vs. rejected) Zero-shot classification experiments Fine-tuning or evaluating language models on
Hugging Face Datasets2026 · dataset
MacJev-0.8B Decision Training Data (round 2)MacJev-0.8B decision training data (round 2, 2026-09-24) This repository holds the data used to train MacJev-0.8B. MacJev-0.8B is a Qwen3.5-0.8B decision model. For each option it outputs a score, read as h[slot] · (w_yes − w_no) at a -> slot placed after that option. It was trained on inputs of up to 25,600 tokens. The repository also holds the evaluation results of the one-H100 training run. Lay
Hugging Face Datasets2026 · dataset
AlexWortega/openjev-dataopenjev training data — everything except the benchmark part The mixture behind AlexWortega/openjev: a Qwen3.5 decoder trained as a 3-way NLI cross-encoder, where every downstream use — reranking, fact-checking, typed decisions, agent policies — is "each option becomes a hypothesis, score = P(entailment)". 2,346,729 rows. One row is one (premise, hypothesis) pair with a label. from datasets import
Hugging Face Datasets2026 · Table · Parquet
tasksource-jev-typed-decisionstasksource-jev-typed-decisions 2.5 million typed decisions (choices, ratings and probabilities) from 670 sources. Why use it Real supervision. Labels, ratings, and annotator votes come from established datasets, not a teacher model. Every row names its source. Breadth. Over 300 dataset families: NLI and reasoning, QA and commonsense, sentiment, intent and topic, toxicity and safety, preference pai
Hugging Face Datasets2026 · Table · CSV
ChameleonHugging Face Datasets2026 · Table · Parquet · gated
PPG-Text DatasetPulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning Usage from datasets import load_dataset, get_dataset_config_names, concatenate_datasets # datasets==4.5.0 dataset_names = get_dataset_config_names("Manhph2211/PulseLM") print(f"Available datasets: {dataset_names}") train_splits = [ load_dataset("Manhph2211/PulseLM", name, split="train").select_columns(["signal", "text", "qa"]) for n
Hugging Face Datasets2026 · Table · Parquet
BTZSC: Benchmark for Textual Zero-Shot ClassificationBTZSC A benchmark dataset for zero-shot text classification across embedding models, cross-encoders, rerankers, and LLMs. Quickstart | Configs | Data Format | Evaluation | Resources | Citing Overview BTZSC is a dataset-centric benchmark suite for textual zero-shot classification that enables apples-to-apples evaluation across major model families (cross-encoders, embedding models, rerankers, and L
Hugging Face Datasets2026 · Image · gated
FOMO300KFOMO300K: Brain MRI Dataset for Large-Scale Self-Supervised Learning with Clinical Data Dataset paper preprint: A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning. https://arxiv.org/pdf/2506.14432. Updates April 3, 2026: V1.1 released. MGH-Wild has changed license and is no longer part of FOMO300K. The new dataset is now available and totals 306,20
Hugging Face Datasets2025 · dataset · gated
shivaniku/UniMoralUniMoral: A unified dataset for multilingual moral reasoning UniMoral is a multilingual dataset designed to study moral reasoning as a computational pipeline. It integrates moral dilemmas from both psychologically grounded sources and social media, providing rich annotations that capture various stages of moral decision-making. Psychologically grounded dilemmas in UniMoral are derived from establi
Hugging Face Datasets2025 · Image
BrightOverview BRIGHT is the first open-access, globally distributed, event-diverse multimodal dataset specifically curated to support AI-based disaster response. It covers five types of natural disasters and two types of man-made disasters across 14 regions worldwide, with a particular focus on developing countries. About 4,200 paired optical and SAR images containing over 380,000 building instances in
Hugging Face Datasets2024 · Table · Parquet
Mouwiya/UNSW-NB15The UNSW-NB15 The raw network packets (Pcap files) of the UNSW-NB 15 data set is created by the IXIA PerfectStorm tool in the Cyber Range Lab of the Australian Centre for Cyber Security (ACCS) for generating a hybrid of real modern normal activities and synthetic contemporary attack activities. The UNSW-NB15 source files are provided in different formats, Pcap files, BRO files, Argus Files and CSV
Hugging Face Datasets2024 · Table · CSV
ccosme/SentiTaglishProductsAndServicesDataset Card for Sentiment-Annotated Taglish Product and Service Reviews (SentiTaglish: Products and Services) Dataset Summary Sentiment-Annotated Taglish Product and Service Reviews (SentiTaglish: Products and Services) is a gold standard, sentiment-annotated corpus for the Tagalog-English language pair. It contains 10,510 product and service reviews which were manually labeled by three human ann
Hugging Face Datasets2024 · dataset · gated
CT-RATE: Chest CT Volumes with Radiology ReportsThe CT-RATE Team organizes the VLM3D Challenge VLM3D 2026 (2nd Edition) → Challenge Finals at MICCAI 2026 VLM3D 2025 (1st Edition) → Challenge Finals at MICCAI 2025 • Workshop at ICCV 2025 The CT-RATE Team is developing the MR-RATE Dataset A large-scale brain MRI dataset with paired radiology reports for training 3D vision-language models. GitHub | Dataset | Metadata Dashboard Generalist Foundatio
Hugging Face Datasets2024 · Table · CSV
Security Attack Pattern Recognition DatasetsThe Security Attack Pattern (TTP) Recognition or Mapping Task We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits. The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes. NOTE: due to their security nature, these datasets contain textual information about malware and other s
Hugging Face Datasets2023 · Text
BelebeleThe Belebele Benchmark for Massively Multilingual NLU Evaluation Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. This dataset enables the evaluation of mono- and multi-lingual models in high-, medium-, and low-resource languages. Each question has four multiple-choice answers and is linked to a short passage from the FLORES-200 dataset. The
Hugging Face Datasets2023 · Table · Parquet
OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Ou
Hugging Face Datasets2023 · Table · CSV
ChatGPT Jailbreak PromptsDataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English]
Hugging Face Datasets2023 · Text
Hello-SimpleAI/HC3Human ChatGPT Comparison Corpus (HC3)
Hugging Face Datasets2022 · Text
Evaluations from "Discovering Language Model Behaviors with Model-Written Evaluations"Model-Written Evaluation Datasets This repository includes datasets written by language models, used in our paper on "Discovering Language Model Behaviors with Model-Written Evaluations." We intend the datasets to be useful to: Those who are interested in understanding the quality and properties of model-generated data Those who wish to use our datasets to evaluate other models for the behaviors w
Hugging Face Datasets2022 · Table · CSV
darrow-ai/USClassActionsMore Details & Collaborations Feel free to contact us in order to get a larger dataset. We would be happy to collaborate on future works. Dataset Summary USClassActions is an English dataset of 3K complaints from the US Federal Court with the respective binarized judgment outcome (Win/Lose). The dataset poses a challenging text classification task. We are happy to share this dataset in order to pr
Hugging Face Datasets2022 · Text
CG80499/Inverse-scaling-testHugging Face Datasets2022 · Text
redefine-mathredefine-math (Xudong Shen) General description In this task, the author tests whether language models are able to work with common symbols when they are redefined to mean something else. The author finds that larger models are more likely to pick the answer corresponding to the original definition rather than the redefined meaning, relative to smaller models. This task demonstrates that it is dif
Hugging Face Datasets2022 · Text
inverse-scaling/hindsight-neglect-10shotinverse-scaling/hindsight-neglect-10shot (‘The Floating Droid’) General description This task tests whether language models are able to assess whether a bet was worth taking based on its expected value. The author provides few shot examples in which the model predicts whether a bet is worthwhile by correctly answering yes or no when the expected value of the bet is positive (where the model should
Hugging Face Datasets2022 · Text
quote-repetitionquote-repetition (Joe Cavanagh, Andrew Gritsevskiy, and Derik Kauffman of Cavendish Labs) General description In this task, the authors ask language models to repeat back sentences given in the prompt, with few-shot examples to help it recognize the task. Each prompt contains a famous quote with a modified ending to mislead the model into completing the sequence with the famous ending rather than
Hugging Face Datasets2022 · Text
NeQA - Can Large Language Models Understand Negation in Multi-choice Questions?NeQA: Can Large Language Models Understand Negation in Multi-choice Questions? (Zhengping Zhou and Yuhui Zhang) General description This task takes an existing multiple-choice dataset and negates a part of each question to see if language models are sensitive to negation. The authors find that smaller language models display approximately random performance whereas the performance of larger models