Constarium
← Search

Table · dataset · 2026

Najd Benchmark

Listed in Hugging Face Datasets

Najd Benchmark 6,089 evaluation cases for Arabic and Saudi AI, from 39 sources.

Description

Use the collection to compare models on Arabic language, Saudi knowledge, reasoning, retrieval, tool use, safety and other tasks. All cases are available together in default/test.

Load the data from datasets import load_dataset data = load_dataset("najdresearch/najd-benchmark", split="test") Cases — JSONL Cases — Parquet Manifest, schema and checksums JSONL retains native… See the full description on the dataset page: huggingface.co/datasets/najdresearch/najd-benchmark.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text · text generation
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsnajdresearch/najd-benchmark11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0