Constarium
← Search

Text · dataset · 2026

CRJudgeBenchmark

Listed in Hugging Face Datasets

CRJudgeBenchmark CRJudgeBenchmark evaluates whether a model can judge the technical trustworthiness of a code-review comment in its pull-request context.

Description

It contains 1,199 labeled examples from 124 pull requests across 9 GitHub repositories, with fixed train, validation, and test splits. Each example pairs a pull request with one target review comment.

It includes the PR description, a base commit, a review patch, linked issues, and a PR activity timeline. The binary trustworthy… See the full description on the dataset page: huggingface.co/datasets/dcloud347/CRJudgeBenchmark.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text · text classification
Provenance · 1 source records, 9 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsdcloud347/CRJudgeBenchmark12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:text-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0