MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Paper • 2608.23035 • Published 13 days ago • 41
Agents-A1 Collection Agents-A1 is a Long-horizon Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. • 12 items • Updated Jul 16 • 42
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models Paper • 2606.16140 • Published Jun 15 • 127
MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery Paper • 2606.06473 • Published Jun 4 • 22
ACC: Compiling Agent Trajectories for Long-Context Training Paper • 2605.21850 • Published May 21 • 62