Data · dataset · 2026
The Rewards of Reinforcement Learning
Listed in ScienceDB
This is the supporting raw training dataset for the paper "An Offline Experience Purification Framework for Suboptimal Expert-Guided Reinforcement Learning with UAV Tracking Application".
Description
It records the cumulative reward per episode during the reinforcement learning training of 6 ablation algorithms on the UAV ground target tracking task, which can fully reproduce the learning curves and convergence speed evaluation results in the paper.Included Algorithm VariantsThe dataset covers all control and ablation algorithms in the paper:Baseline (Pure SAC): Classical vanilla Soft Actor-Critic algorithm without pre-training or experience injection, serving as the performance benchmark.No Filter: SAC with actor-only behavioral cloning (BC) pre-training using unfiltered raw suboptimal expert data.AC Pre-train: SAC with joint actor-critic pre-training using filtered suboptimal expert data.Ours: The proposed method — SAC with actor-only BC pre-training based on filtered suboptimal expert data.No BC: SAC variant without BC pre-training, with only filtered suboptimal experience injected into the replay buffer.No Replay: Variant of the proposed method with the experience replay mechanism removed.Data SpecificationAll algorithms are trained under identical environment parameters, network architectures and experimental protocols, with a maximum of 3000 episodes per training run.Each algorithm contains complete records from 5 to 15 independent repeated training runs to support statistical significance analysis.The data is stored as per-episode cumulative rewards, which can be directly used for plotting learning curves, calculating convergence metrics, and conducting algorithm comparison studies.ApplicationsThis dataset can reproduce all experimental conclusions of the multi-metric convergence speed evaluation system in the paper.
It can also serve as a benchmark dataset for performance verification of other demonstration-guided reinforcement learning and convergence acceleration methods.
Links
Where it is published
- DOI doi.org/10.57760/sciencedb.43249 ↗
DOI / persistent id · from scidb cn
Catalogue records · 1
- OAI-PMH record scidb.cn/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=10.57760%2… ↗
metadata API · from scidb cn
Topics
- From keywords
- Earth & Environmental Science · Engineering · Humanities · Life Sciences · Social Science
Provenance · 1 source records, 9 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| ScienceDB | 10.57760/sciencedb.43249 | 8 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| concepts[field].local:field:earth-environmental | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:engineering | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:humanities | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:social-science | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| description | source · scidb cn | connector:scidb_cn@1.0.0 | /metadata/dc/description |
| license_text | source · scidb cn | connector:scidb_cn@1.0.0 | |
| publication_date | source · scidb cn | connector:scidb_cn@1.0.0 | |
| title | source · scidb cn | connector:scidb_cn@1.0.0 | /metadata/dc/title |