Constarium
← Search

Table · dataset · 2026

llm-jp/llm-jp-4.1-8b-thinking-dpo-data

Listed in Hugging Face Datasets

llm-jp-4.1-8b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4.1-8b-thinking.

Description

It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation.

The fields chosen_analysis… See the full description on the dataset page: huggingface.co/datasets/llm-jp/llm-jp-4.1-8b-thinking-dpo-data.

Links

Topics

Stated by source
text
Provenance · 1 source records, 8 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsllm-jp/llm-jp-4.1-8b-thinking-dpo-data12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0