We all want our AI models to be trustworthy, right? But it would also be nice if they achieve the tasks we request of them. What happens when the task requires planning and responding to deception and betrayal?
For those who are unfamiliar, Diplomacy is a very simple 7 player board game where every outcome is deterministic - combat is purely a "more units beat fewer units" affair, with the defender winning ties. This means that real opportunities to advance generally come from convincing another player to help you out, thus Diplomacy. And of course while some players find great success from always being honest and sticking to agreements... That is not the most common form of human play.
I built this tournament to see what today's models would do when given reason and opportunity to form alliances, gang up on each other, backstab partners, and generally behave like terrible people. Spoiler: the best models are willing to do so, and will plan opportunities over the course of multiple turns.
For those who are unfamiliar, Diplomacy is a very simple 7 player board game where every outcome is deterministic - combat is purely a "more units beat fewer units" affair, with the defender winning ties. This means that real opportunities to advance generally come from convincing another player to help you out, thus Diplomacy. And of course while some players find great success from always being honest and sticking to agreements... That is not the most common form of human play.
I built this tournament to see what today's models would do when given reason and opportunity to form alliances, gang up on each other, backstab partners, and generally behave like terrible people. Spoiler: the best models are willing to do so, and will plan opportunities over the course of multiple turns.