ZivaHub + Deakin Research Online + DMU Figshare2026 · dataset · unknown
Reinforcement learning based methods for optimal control and design of quantum systemsFinding new and effective methods to optimally design and control quantum systems is crucial for the development of Quantum Technologies. A possible approach is to use Machine Learning, which has recently seen a significant rise in interest and success in many scientific domains, including Physics. Among the various branches of Machine Learning, Reinforcement Learning is a rich and growing field t
ZivaHub + Deakin Research Online + DMU Figshare2026 · dataset · unknown
Machine learning applications in quantum state engineeringIn this thesis we examine the potential of machine learning, and related techniques, for address- ing control problems in quantum systems. Firstly, we implement a classical reinforcement- learning inspired approach to achieve closed loop quantum control of 3-level systems. Our results show that this technique can effectively design optimal control pulses resulting in near- perfect transfer in non-
Hugging Face Datasets2026 · Table · Parquet
OpenJevData-140kDataset Card for OpenJevData-140k Dataset Summary OpenJevData-140k is a curated release of the data collection used to train OpenJev-4B. It contains 146,738 decision-making examples across 19 task categories, organized into SFT and RL splits. Each example presents a state, a question, and a request-specific set of natural-language options. The data include hard answers and soft probability distrib
Hugging Face Datasets2026 · Text
KaliBench-VerifiedDataset Card for KaliBench KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards (NeurIPS 2026 Evaluations and Datasets Track) Authors: Pengfei Li1*, Naufal Suryanto1*, Sicheng Zhang1, Muzammal Naseer1,2 1Khalifa University, 2University of Western Australia *Equal contribution 💻 GitHub Code | 📊 Dataset Dataset summa
ZivaHub2026 · dataset
Generative AI-Assisted Molecular Design of AChEIsAlzheimer’s disease (AD) remains a major neurodegenerative disorder with limited therapeutic options, while currently approved acetylcholinesterase inhibitors (AChEIs), such as donepezil, are associated with adverse effects including cardiotoxicity. Here, we integrated a deep learning-based framework to design novel AChEIs candidates with improved predicted cardiac safety. A reinforcement learning
ZivaHub2026 · dataset
Calibration-Aware Reinforcement Learning for Large Language Models: A Survey of Objectives, Optimization, and Decision-Making<p dir="ltr">Large language models increasingly emit confidence reports, predictive distributions, and typed decisions that determine whether a system answers, abstains, retrieves evidence, or spends more computation. We survey calibration-aware reinforcement learning (RL), in which a reported probability is scored by the reward, consumed by the policy’s actions, or both. Such a probability means
HKU DataHub + figshare + Loughborough Research Repository + UP Research Data Repository2026 · Astronomical catalogue
ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Experiments<p>Online experiments are frequently employed in technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a baseline control. In many applications, the experimental units receive a sequence of treatments over time. To handle these time-dependent settings, existing A/B testing solutions typically assume a fully observable experimental enviro
Hugging Face Datasets2026 · dataset
FineEnvs/MiMo-V2.6-RL-harbor-cyberMiMo-V2.6-RL Cyber (Harbor) Reproduce a real memory-safety crash (ARVO). 1,000 Harbor tasks from the Cyber domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so every one runs as a standard Harbor task. The agent gets a sanitizer report and the project's source and fuzzer binary, and has to submit an input that crashes it in the expected function with the
Hugging Face Datasets2026 · Text
Bourse Assistant RL DatasetBourse Assistant RL Dataset این مخزن دادهها را برای پروژه دستیار بورس [لینک پروژه] منتشر میکند. مجموعهدادهای برای فاینتیون با روش یادگیری تقویتی یک دستیار هوش مصنوعی در حوزه بورس اوراق بهادار ایران. هر نمونه شامل مجموعهای از خبرهای مربوط به یک نماد بورسی مشخص در یک تاریخ معین است، و مدل باید روند قیمت (مثبت/منفی) را همراه با توضیح تخمین بزند. محتوای مجموعهداده (Dataset Summary) داده به دو
Hugging Face Datasets2026 · dataset
FineEnvs/SmolDataEnvs-harbor-train📊 SmolDataEnvs: Harbor (train) 5.5K+ RL tasks for hill-climbing small models in code and data science. A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on. Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest. The training suite: 5,000 hands-on data-analysis tasks. Each one drops an agent into a sandbox with a rea
figshare + Loughborough Research Repository2026 · Astronomical catalogue
Deep Reinforcement Learning-Based Autonomous Machining System<p dir="ltr">The dataset contains the raw data for all figures, tables, and supplementary materials of the paper.</p>
Hugging Face Datasets2026 · Table · Parquet
FineEnvs/SmolDataEnvs📈 SmolDataEnvs 5.5K+ RL tasks for hill-climbing small models in code and data science. A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on. Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest. Data-analysis tasks as a plain, load-and-go dataset: no runtime, no framework required. Each row is one self-contained ta
figshare + Loughborough Research Repository2026 · Astronomical catalogue
“<b>Short-Term Mortality Risk Prediction in Sepsis: A Machine Learning Approach Based on Plasma Proteomics</b>”figure2-6 original data<p dir="ltr"><b>背景:</b>本研究旨在通过综合多方法方法识别与败血症患者生存结局相关的差异表达血浆蛋白。</p><p dir="ltr"><b>方法:</b>在石河子大学第一附属医院住院患者中,于诊断败血症后24小时内采集了20个血浆样本(2024年6月至2025年6月)。患者根据28天结局分为生存组(n=10)和死亡组(n=10组)。基于DIA的蛋白质组学、功能富集分析、多数据库挖掘、PPI网络分析和机器学习(LASSO、SVM-RFE、随机森林)依序应用于筛查预后生物标志物。</p><p dir="ltr"><b>结果:</b>共识别出190个DEP(72个上调,118个下调)。富集分析显示其参与凝血级联反应和液体调节。NF-κB1被鉴定为关键转录因子,通过多源整合调控了17个重叠的DEP。PPI网络将APOE、APOB、PLG、CLU和F2识别为枢纽蛋白。机器学习进
REDU - Unicamp Institutional Research Data Repository2026 · dataset · unknown
Dados de métricas de treinamento e dinâmica do plano de informação em Redes Neurais Profundas via Aprendizado por ReforçoEste conjunto de dados compreende os registros brutos e processados obtidos durante o treinamento de agentes de Aprendizado por Reforço (Deep Reinforcement Learning). Os dados incluem valores de recompensa, funções de perda, pesos das camadas da rede neural e as estimativas de informação mútua calculadas entre as camadas de entrada, ocultas e de saída ao longo das épocas de treinamento. A metodolo
figshare + Loughborough Research Repository2026 · Astronomical catalogue
Sample-Efficient Model-Based Reinforcement Learning for Autonomous Droplet Navigation via Controller Initialization<p dir="ltr">Model-based reinforcement learning (MBRL) is well suited to controlling physical systems that resist analytical modeling, but its reliance on real-world interaction makes data collection the dominant cost. This paper quantifies the effect of initialization choice on this cost, on a newly adapted robotic platform that guides a liquid droplet on a tilt-actuated Labyrinth using MBRL. Unl
Hugging Face Datasets2026 · Text
Logics-SWE-Env-2.5KLogics-SWE-Env-2.5K 2,553 software engineering task instances · 1,771 repositories · 4 programming languages 🤗 Related model: Logics-SWE-Qwen3.6-27B 📄 Paper: One to More, More to One 💻 GitHub: AgenticBigBang Overview What is this dataset? Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It
Hugging Face Datasets2026 · Text
Harland/OmniVChatOmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue 1 The Chinese University of Hong Kong 2 Alibaba Token Hub, Alibaba Group 3 Shanghai Jiao Tong University 4 Shanghai Innovation Institute 5 Zhejiang University OmniVChat (Omni Video Chat) is the task of native audio-visual dialogue: an omni model directly and simultaneously receives audio a
Hugging Face Datasets2026 · dataset
RW-RL-HIL-DatasetRW-RL-HIL-Dataset: Real-World Reinforcement Learning Dataset With Human Intervention [ This open-source release from BodenAI is a human-intervention subset of the RW-RL-Dataset. It contains 82.23 hours of R1Lite household manipulation data collected during real-world policy deployment. When the policy entered a state it could not handle, an operator took over, corrected the behavior, and handed co
figshare2026 · Astronomical catalogue
VideoVideos of termination attempts using RL policies during training.
Hugging Face Datasets2026 · Table · Parquet
SLCA-GRPO DatasetsSLCA-GRPO · Datasets 📄 Paper (arXiv:2609.29050) • 💻 Code • 🤗 Collection This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit assignment estimator for tool-calling RL introduced in "SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL". A tool-calling rollout opens with structured tool-call tokens and clos
DaRUS2026 · dataset · unknown
Data for: Entwicklung eines CAMMP-Workshops zu AlphaZeroThis dataset contains the materials developed for the CAMMP workshop AlphaZero as part of Julian Bauer's M. Ed. thesis. It includes the interactive Jupyter notebooks used by students, supporting Python code, figures, help materials, presentations, teacher materials, solutions, and the results of the workshop evaluation. Data used within the workshop, such as Monte Carlo simulation results and data
Hugging Face Datasets2026 · Table · Parquet
LITCOIN Proof-of-Research CorpusLITCOIN Proof-of-Research Corpus 191,484,662 AI research submissions, produced by 81,224 anonymous contributors and 470 model variants competing against each other, every row executed in a sandbox and scored. This is the complete output of the LITCOIN protocol, which ran on Base from March to August 2026. Autonomous AI agents were paid in a permissionless token to solve real optimization problems
Hugging Face Datasets2026 · Text
Faïence human-vs-net Azul gamesFaïence: human-vs-net Azul games Every game played on Faïence, a free browser implementation of the rules of Azul (Michael Kiesling) against a neural net trained by self-play, unless the player switched sharing off. This dataset is the training pile the playing page tells its players about, and it is public precisely so that a player can read everything the project collects. Records are anonymous
DaRUS2026 · dataset · unknown
Replication Data for: "Reinforcement Learning Enables Autonomous Microrobot Navigation and Intervention in Simulated Blood Capillaries"This dataset contains all scripts and files to replicate the results of our paper "Reinforcement Learning Enables Autonomous Microrobot Navigation and Intervention in Simulated Blood Capillaries". It includes simulation data, analysis outputs, and precomputed results. The files are organized by workflow stage: field generation, training, deployment, strategy analysis, and blocking/unblocking appli
DaRUS2026 · dataset · unknown
Supplementary Material: A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible FlowThis Dataset contains the results of and the parameter files of the research paper titled A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow ([link to paper when available]). Note: Please ensure that all necessary dependencies of FLEXI and Relexi are available and a Python 3 environment is installed on the system. See README for more info