Constarium
← Search

Table · dataset · 2025

SWE-bench Pro

Listed in Hugging Face Datasets

Description

SWE-bench Pro V2 SWE-bench Pro is a benchmark of long-horizon software engineering tasks drawn from real pull requests in 11 open-source repositories (Go, Python, JavaScript, TypeScript). Each task gives an agent a repository at a base commit plus a PR description with explicit requirements and interfaces; the agent's patch is graded by hidden fail-to-pass and pass-to-pass tests in a pristine container. V2 (2026-09-22) is the default config: 642 tasks (go 256, python 237, js 145… See the full description on the dataset page: huggingface.co/datasets/ScaleAI/SWE-bench_Pro.

Links

Where it is published

Documentation and papers

Catalogue records · 1

Topics

Stated by source
text
Provenance · 1 source records, 8 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsScaleAI/SWE-bench_Pro12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0