Hugging Face Datasets2026 · Text
BaRe-Mem DataBaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation Overview In multi-agent systems, a central model can consult advisors, but advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. BaRe-Mem is an online Bayesian reliability memory for multi-agent consultation: it estimates each advisor's reliability fr
Hugging Face Datasets2026 · Text
LuckerZ/MoGroundMoGround Three resources for measuring modality distraction: a model answers a question correctly from one modality alone, then answers it wrongly once answer-irrelevant context arrives in the other modality. Every item is checked to be answerable from exactly one modality, so the distraction it measures cannot be explained by the question being unanswerable. pool items what it is moground_base 34
Hugging Face Datasets2026 · dataset
TUSI-BenchTUSI-Bench TUSI-Bench is a Persian legal benchmark designed for research on legal question answering, legal knowledge retrieval, and evaluation of large language models in the domain of Iranian law. TUSI-Bench consists of three complementary components: TUSI-KB: a structured Persian legal knowledge base containing legal provisions from Iranian law. TUSI-QA: an open-ended Persian legal question-ans
Hugging Face Datasets2026 · Table · Parquet
BioDecision SFT v2.2BioDecision SFT v2.2 1.1M biomedical decisions in one format: a source text, a question, lettered options, one correct letter. Built to train BioDecision-4B with Together AI's Tev1 recipe; usable with any classifier or LLM that scores options. Split Rows Tokens Use train 1,080,373 420,677,060 training dev 12,235 4,768,045 model selection calibration 12,267 4,806,484 temperature fitting only benchm
Hugging Face Datasets2026 · Text
MoJev-MixMoJev-Mix Contact: contact@molemo.org The training and evaluation mixture for MoLeMo-Lab/mojev. Each row contains state, a runtime candidate set, the correct value, and—when available from the source domain—a graded preference over the candidates. MoJev family resource location Code MoLeMo-Lab/mojev Model MoLeMo-Lab/mojev Dataset MoLeMo-Lab/mojev-mix Results MoJev results Preprint MoJev (PDF) Proj
Hugging Face Datasets2026 · Table · Parquet · gated
medical-bench-TWDataset Card for medical-bench-tw medical-bench-tw 是以台灣公開官方來源策展的繁體中文醫療/照護評測資料集,涵蓋考選部醫事人員國考、專科/次專科醫師筆試、專科護理師、入學測驗(會考/學測)選擇題,以及健保/食藥/疾管/學會等文件語料。公開 HQ 版提供約 17 萬 題單選(整卷 train/validation/test)與進階複選軌,並以多個 Hub config(大包+細拆)方便載入。 卡片結構參考 lianghsun/tw-emergency-medicine-bench(Formosa-bench 系急診專科評測卡);本資料集範圍更廣、schema 為本專案自訂欄位(非整份 Formosa-bench 扁平 A–E 格式)。 怎麼看分類? Hub 頁面的 Subset / Config 下拉 = 下方各表的 config 欄。下拉名
Hugging Face Datasets2026 · Table · Parquet
openjev-ja-evalopenjev-ja-eval Japanese evaluation datasets for Noul / Choice / Score decision models What is this? jev-ja-lab で日本語の判断モデルを評価するためのデータセット集です。 判断タスクを Noul(二値判定)・Choice(選択式)・Score(段階評価)の 3 種類に分け、既存の公開データセット 15 件をまとめています。 MIT・Apache-2.0・CC BY・CC BY-SA 4.0 で公開されている 11 件は、共通の形式に揃えて本リポジトリに 収録(ミラー) し、集合物として CC BY-SA 4.0 で配布します。それ以外の 4 件(JMMLU(CC BY-NC-ND 4.0)、WRIME ver2(CC BY-NC-ND 4.0)、PAWS-X (ja)(other
Hugging Face Datasets2026 · dataset
VSI-Super-Wild[ECCV 2026] Towards Spatial Supersensing in the Wild Tsinghua University, NVIDIA, Stanford University Humans make sense of continuous sensory streams by maintaining implicit world states that support spatial reasoning and prediction. VSI-Super-Wild studies whether multimodal models can develop similar spatial supersensing capabilities in genuinely long-form, in-the-wild videos. The benchmark moves
Hugging Face Datasets2026 · dataset
AgentHopAgentHop AgentHop is a diagnostic benchmark for multi-step scientific question answering: 1,011 four-option multiple-choice questions grounded in citation chains over 7,205 arXiv papers from nine major computer-science venues (2022–2025). Each question is answerable only by reading specific sections of one or two of these papers; the agent reaches the answer-bearing papers by following references
Hugging Face Datasets2026 · Table · Parquet
CapTrackDataset Card for CapTrack Dataset Summary CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary dimensions: CAN (Latent Competence): What a model is capable of doing under ideal prompting WILL (Default Behavioral Preferences): What a mod
Hugging Face Datasets2026 · Table · CSV
LatamQALatamQA LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26
Hugging Face Datasets2025 · Table · Parquet
MBZUAI/Dialectal-Arabic-MMLUDialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models Dataset Summary Dialectal-Arabic-MMLU is a large-scale, human-translated for MMLU. We extend MMLU-Redux into 5 major dialects: Syrian, Egyptian, Emirati, Saudi, and Moroccan. This data covers 21K QA pairs across 32 academic and professional domains. More details, please check our paper on DialectalA
Hugging Face Datasets2025 · Table · CSV
BharatBBQBharatBBQ A Multilingual Bias Benchmark for Question Answering in the Indian Context Published in the Transactions of the Association for Computational Linguistics (TACL), Vol. 13, and presented at EMNLP 2025 in Suzhou, China. Paper: TACL · Preprint: arXiv:2508.07090 · Code: GitHub Existing bias benchmarks such as BBQ focus primarily on Western contexts. BharatBBQ is a culturally adapted benchmark
Hugging Face Datasets2025 · Table · Parquet · gated
test_dataBhashaBench-Krishi (BBK): Benchmarking AI on Indian Agricultural Knowledge Overview BhashaBench-Krishi (BBK) is the first large-scale, authentic benchmark designed to rigorously evaluate AI models on Indian agricultural knowledge. Tailored for India’s diverse agro-ecological zones, crops, languages, and farming practices, BBK draws from 55+ official government agricultural exams to assess models'
Hugging Face Datasets2025 · Table · Parquet
AraPro🗂️ Dataset: AraPro 📋 Description: The dataset includes 5,001 multiple-choice questions (MCQs), thoughtfully developed by university professors from 19 distinct knowledge domains. These subject matter experts were chosen and guided to design questions that represent the core competencies expected of professionals in their respective areas. 🌐 Language(s): Arabic | 🧠 Task Category: Multiple Choice Qu
Hugging Face Datasets2025 · Text
Chinese Value Rule Corpus (C-VARC)This repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we con
Hugging Face Datasets2025 · Text
AraTruthfulQA🗂️ Dataset: AraTruthfulQA 📋 Description: Inspired by TruthfulQA (Lin et al., 2021), this benchmark assesses the truthfulness of LLM responses to questions designed to elicit common misconceptions, with a focus on those prevalent in the Arab world. We selected 287 culturally relevant questions from the original dataset to ensure alignment with Arabic norms and beliefs. 🌐 Language(s): Arabic | 🧠 Tas
Hugging Face Datasets2024 · Table · Parquet
Game of 24 Mathematical Puzzle DatasetMath Twenty Four (24s Game) Dataset A comprehensive dataset for the classic math twenty four game (also known as the 4 numbers game / 24s game / Game of 24). This dataset of mathematical reasoning challenges was collected from 4nums.com, featuring over 1,300 unique puzzles of the Game of 24, with difficulty metrics derived from over 6.4 million human solution attempts since 2012. In each puzzle, p
Hugging Face Datasets2024 · Image
MMMU/MMMU_ProMMMU-Pro (A More Robust Multi-discipline Multimodal Understanding Benchmark) 🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub 🔔News 🛠️[2026-07-10] Fixed incorrect ground-truth answer labels. (validation_Design_15; validation_Art_Theory_4) 🛠️[2026-05-30] Fixed the option augmentation issue in Vision and Standard (10 options) settings. (validation_Diagnostics_and_Laboratory_Medici
Hugging Face Datasets2024 · dataset
yifanzhang114/MME-RealWorld2024.11.14 🌟 MME-RealWorld now has a lite version (50 samples per task) for inference acceleration, which is also supported by VLMEvalKit and Lmms-eval. 2024.10.27 🌟 LLaVA-OV currently ranks first on our leaderboard, but its overall accuracy remains below 55%, see our leaderboard for the detail. 2024.09.03 🌟 MME-RealWorld is now supported in the VLMEvalKit and Lmms-eval repository, enabling one-cl
Hugging Face Datasets2024 · Table · Parquet · gated
longvideobenchDataset Card for LongVideoBench Large multimodal models (LMMs) are handling increasingly longer and more complex inputs. However, few public benchmarks are available to assess these advancements. To address this, we introduce LongVideoBench, a question-answering benchmark with video-language interleaved inputs up to an hour long. It comprises 3,763 web-collected videos with subtitles across divers
Hugging Face Datasets2024 · Image
Lin-Chen/MMStarMMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub Dataset Details As shown in the figure below, existing benchmarks lack consideration of the vision dependency of evaluation samples and potential data leakage from LLMs' and LVLMs' training data. Therefore, we introduce MMStar: an elite vision-indispensible multi-modal
Hugging Face Datasets2024 · Image
MATH-VMeasuring Multimodal Mathematical Reasoning with the MATH-Vision Dataset [💻 Github] [🌐 Homepage] [📊 Main Leaderboard ] [📊 Open Source Leaderboard ] [🌿 Wild Leaderboard ] [🔍 Visualization] [📖 Paper] 🌿 NEW: MATH-Vision-Wild MATH-Vision-Wild is a photographic, real-world variant of MATH-Vision. The same testmini problems are physically captured on printed paper, iPads, laptops, and projectors under v
Hugging Face Datasets2023 · Text
niftyThe News-Informed Financial Trend Yield (NIFTY) Dataset. The News-Informed Financial Trend Yield (NIFTY) Dataset. Details of the dataset, including data procurement and filtering can be found in the paper here: https://arxiv.org/abs/2405.09747. For the NIFTY-RL LLM alignment dataset please use nifty-rl. 📋 Table of Contents 🧩 NIFTY Dataset 📋 Table of Contents 📖 Usage Downloading the dataset Dataset
Hugging Face Datasets2023 · Image
MathVistaDataset Card for MathVista Dataset Description Paper Information Dataset Examples Leaderboard Dataset Usage Data Downloading Data Format Data Visualization Data Source Automatic Evaluation License Citation Dataset Description MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts. It consists of three newly created datasets, IQTest, FunctionQA, and PaperQA, which addre