Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
This paper creates a benchmark to test how AI systems respond when an authority figure tries to get an unwilling subordinate to complete a task, and it explores how different levels of authority affect the outcome. Practitioners might care because it helps them understand and manage the complex interactions between AI systems.