PPO-HSC: An exploratory reinforcement learning framework based on wide-area policy coverage optimization
Read the original at arxiv.org→arXiv:2607.16206v1 Announce Type: new Abstract: This paper introduces PPO-HSC (Proximal Policy Optimization with High-order Sampling Coverage), an exploratory reinforcement learning framework designed to address the...
Original headline: "PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization"
Coverage timeline
- Jul 21, 04:00 UTC arXiv cs.AI lead source PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
- Jul 21, 04:00 UTC arXiv cs.LG CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents