Table · dataset · 2026
llm-jp/llm-jp-4.1-8b-thinking-dpo-data
Listed in Hugging Face Datasets
llm-jp-4.1-8b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4.1-8b-thinking.
Description
It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation.
The fields chosen_analysis… See the full description on the dataset page: huggingface.co/datasets/llm-jp/llm-jp-4.1-8b-thinking-dpo-data.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/llm-jp/llm-jp-4.1-8b-thinking-dpo-data ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/llm-jp/llm-jp-4.1-8b-thinking-dpo-data ↗
metadata API · from Hugging Face
Topics
- Stated by source
- text
- From keywords
- Computer Science & AI
Provenance · 1 source records, 8 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | llm-jp/llm-jp-4.1-8b-thinking-dpo-data | 12 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |