Constarium
← Search

Table · dataset · 2026

SecondState FAB — Agent Traces and Grading

Listed in Hugging Face Datasets

FAB — Agent Traces and Grading 600 completed agent runs: four models × 50 tasks × three trials.

Description

Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations.

Tasks 041 and 049 are included. The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.… See the full description on the dataset page: huggingface.co/datasets/secondstate/finance-agents-benchmark-traces.

Links

Topics

Stated by source
question answering · tabular · text
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetssecondstate/finance-agents-benchmark-traces7 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:tabularsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:question-answeringsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0