WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 20 days ago • 143
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 25 days ago • 275
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 28 days ago • 151
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 28 days ago • 340
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published Jul 14 • 109
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems Paper • 2606.18837 • Published Jun 17 • 59