HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals Paper • 2609.04444 • Published 19 days ago • 4
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Paper • 2607.15434 • Published Jul 20 • 5
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models Paper • 2606.18142 • Published Jun 17 • 2