🤯 When AI Breaks Its Promises: What a Diplomacy Benchmark Reveals About Agentic AI
A model that can negotiate, deceive, and outsmart its opponents may be impressive. But should we trust it when it tells us, “I’ve finished the job”? That is the uncomfortable question raised by a recent evaluation from Olam Labs , which tested fron…