ETH Zürich Research Collection2026 · Table · CSV
Marble Offcut Dataset: 150 Segmented Offcuts with Thickness and Shape MetricsBinary footprint masks, measured thicknesses and shape metrics for 150 individual physical marble offcuts. Offcuts were photographed on a robot pick platform, located with ArUco fiducials for metric scale, and segmented with point-prompted SAM2. Each offcut carries its footprint mask at platform level (parallax-corrected, rotated so its longest hull edge is axis-aligned, cropped tight), its thickn
ETH Zürich Research Collection2026 · Table · CSV
Leaf-tip keypoint annotations and temporally aligned point cloud sequences of TrackPlant3Ddati.emilia-romagna.it2026 · Table · CSV
Elenco vie comunaliVista delle vie attive gestite dalla toponomastica mediante specifico applicativo nell'ambito del progetto ACI - Dato gestito attraverso l'Anagrafe Comunale degli Immobili (ACI)
Hugging Face Datasets2026 · Text
KaliBench-VerifiedDataset Card for KaliBench KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards (NeurIPS 2026 Evaluations and Datasets Track) Authors: Pengfei Li1*, Naufal Suryanto1*, Sicheng Zhang1, Muzammal Naseer1,2 1Khalifa University, 2University of Western Australia *Equal contribution 💻 GitHub Code | 📊 Dataset Dataset summa
Hugging Face Datasets2026 · Text
Nickyang/RazorCalRazorCal RazorCal is a general-purpose calibration corpus released with RAZOR. RazorCal.json contains 2,048 samples across seven domains (about 8 MB). Each record includes input messages and source attribution. The records carry no architecture-specific fields, so the corpus also suits FFN or layer pruning, perplexity measurement and other calibration-based compression; RAZOR uses it to estimate p
dati.emilia-romagna.it2026 · Table · CSV
Numeri civiciNumeri civici attivi gestiti dalla toponomastica mediante specifico applicativo Anagrafe Comunale degli Immobili (ACI). Il dataset contiene i seguenti campi: ID_CIV (numerico): Identificativo univoco del civico NUMERO (numerico) : Numero civico (esponente escluso) ESPONENTE (varchar): Esponente del numero civico ID_VIA (numerico) : Identificativo unico della Via VIA_NOME_U (varchar) : Nome ufficia
Hugging Face Datasets2026 · Text
GSM8K-UZ (Cyrillic)GSM8K-UZ (Cyrillic) Uzbek GSM8K in the Cyrillic script. Part of a parallel pair: kurbanovxurshidbek/gsm8k-uz-lat and kurbanovxurshidbek/gsm8k-uz-cyr. Both contain exactly the same problems (same idx), differing only in script. Split Rows train 7417 test 1308 Source and construction Latin text: NeuronUz/gsm8k-uz, a machine translation of openai/gsm8k (main) into Uzbek Latin. Cyrillic text: automati
Hugging Face Datasets2026 · Text
HABIT-BenchHABIT-Bench Evaluating Habit Induction from Longitudinal Weak Evidence in Agent Memory HABIT-Bench evaluates whether a memory system can infer a latent habit from repeated, individually incomplete observations and apply it only when the current context supports it. Tasks test weak-evidence induction, applicability boundaries, local exceptions, temporal changes, and the distinction between user-end
Hugging Face Datasets2026 · Text
The First And Best Fable 5.1 Reasoning DataDataset Description This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces
Hugging Face Datasets2026 · Text
PublicHearingLDSExtendedPublicHearingLDSExtended Dataset PublicHearingLDSExtended é um dataset estruturado para verificação de alegações (claim verification) e recuperação de evidências em audiências públicas da Câmara dos Deputados do Brasil. Ele estende o dataset original PublicHearingBR_LDS, preservando 100% dos seus documentos e metadados, adicionando uma camada estruturada de alegações (claims) por orador com evidên
Hugging Face Datasets2026 · Text
GSM8K-UZ (Latin)GSM8K-UZ (Latin) Uzbek GSM8K in the Latin script. Part of a parallel pair: kurbanovxurshidbek/gsm8k-uz-lat and kurbanovxurshidbek/gsm8k-uz-cyr. Both contain exactly the same problems (same idx), differing only in script. Split Rows train 7417 test 1308 Source and construction Latin text: NeuronUz/gsm8k-uz, a machine translation of openai/gsm8k (main) into Uzbek Latin. This dataset reproduces the L
Hugging Face Datasets2026 · Text
Agentic Coding Chain-of-Thought Dataset🤖 Agentic Coding CoT Dataset A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with e
Hugging Face Datasets2026 · Text
lucasfrag/fact-checking-abstention-saes-dataFact-Checking Abstention SAEs — data Companion data for lucasfrag/fact-checking-abstention-saes. Per base model (e.g. llama-3.1-8b-instruct/) file content examples.jsonl the exact ordered list of training claims (VitaminC + FEVER train splits, gold evidence), one JSON per line: claim, evidence (list), label (SUP/REF/NEI), source. Line i is prompt i of the harvest, so activations can be regenerated
Hugging Face Datasets2026 · Text
liuyueyi-8/Elastic-Forcing-training-datasetElastic-Forcing training datasets wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351. Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales. wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,4
Hugging Face Datasets2026 · Text
BaRe-Mem DataBaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation Overview In multi-agent systems, a central model can consult advisors, but advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. BaRe-Mem is an online Bayesian reliability memory for multi-agent consultation: it estimates each advisor's reliability fr
Hugging Face Datasets2026 · Text
EngiWorld Environment ImagesEngiWorld Environment Images This repository contains virtual machine images for the EngiWorld benchmark, covering 15 open-source engineering software environments. You can find more information in our paper, EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments? Paper (arXiv): Coming soon Project website: engiworld.github.io Project GitHub: Hongcheng-Gao/EngiWorld L
Hugging Face Datasets2026 · Text · gated
Meddies ASR BenchMeddies ASR Bench This public dataset repository contains the validated 10 full audio chunks used for the MOSS + Gemini 3.8 Flash smoke benchmark. Contents data/full_10_chunks/audio/: ten full MP3 input chunks. data/full_10_chunks/moss/: the ten source MOSS-Diarize-Transcribe JSON records. data/full_10_chunks/manifest.json: chunk paths, offsets, durations, and embedded MOSS text. results/: the dow
Hugging Face Datasets2026 · Text
HUMMBL 40k Multi-Agent Wicked Problems & Coordination CorpusHUMMBL 40k Multi-Agent Wicked Problems & Coordination Corpus A foundational 40,171-event empirical dataset capturing real-world multi-agent coordination, epistemic problem decomposition, failure mode taxonomies, and strategic intelligence surges generated across the HUMMBL autonomous agent fleet. Dataset Overview The dataset provides structured visibility into how autonomous agents navigate comple
Hugging Face Datasets2026 · Text
Naitik-sudo123/gsm8k-verifier-training-dataHugging Face Datasets2026 · Text
nabil420/dataHugging Face Datasets2026 · Text
LiveChatBenchLiveChatBench (v1.1) Korean → English translation benchmark built from Korean live-streaming chat messages. Each row pairs a Korean chat message with an English reference translation. Where a message relies on streaming slang, community jargon, or a streamer/game name, a short background note explains the term; it is an empty string when no note is needed. Fields Field Type Description background
Hugging Face Datasets2026 · Text
SecondState FAB — Finance Agents BenchmarkFAB — Finance Agents Benchmark FAB is an open-source project for benchmarking LLM agents' ability to perform financial due diligence in a synthetic company data room. FAB consists of a dataset of tasks containing agent instructions, documents and rubrics, and an execution harness for running and evaluating agents. This repository contains the dataset; the harness is available on GitHub. Dataset 50
Hugging Face Datasets2026 · Text
SwarmTraces publisher artifacts and Sev observationsSwarmTraces publisher artifacts Sev-4B v0.3.0 research update The v0.3.0 response-policy checkpoint improves authored policy decisions from 127/175 to 147/175 across 35 held-out source programs. At its calibration-selected alert threshold, it flags 14/140 permitted decisions, a 10% false-alert rate on this panel. 26/28 registered checks pass. The DNS diagnostic remains 21/32, with one repaired ans
Hugging Face Datasets2026 · Text
TozAI/Toz-1-SFTTöz-1 SFT verisi Töz-1 sohbet modelini eğitmek için kullanılan Türkçe SFT (talimat/sohbet) verisi: 12,297 örnek, 14,316 soru-cevap turu. Verinin neredeyse tamamı kodla üretildi (uretim/ klasöründeki scriptler). Amaç modele bilgi değil biçim öğretmek: 195M parametreli bir modele SFT'de bilmediği bir bilgiyi öğretmek uydurmayı öğretir (Gekhman ve ark., 2024). Bu yüzden: Bilgi soruları sadece doğru c
Hugging Face Datasets2026 · Text
ZorQelis Decision Suite (ZDS-1)ZorQelis Decision Suite (ZDS-1) One benchmark for decision models. A decision model reads a situation (the state), a question and a fixed list of options, and returns one option with a calibrated confidence. No free text. ZDS-1 measures that job on 22,864 cases in three tracks, and scores every system with the same harness, the same metrics and the same pricing rule. Cases 22,864 in 3 tracks Metri