Google had its LLM murder itself in Werewolf to test its AI smarts

At GDC 2024, Google AI senior engineers Jane Friedhoff (UX) and Feiyang Chen (Software) showed off the results of their Werewolf AI experiment, in which all the innocent villagers and devious, murdering wolves are Large Language Models (LLMs).

Friedhoff and Chen trained each LLM chatbot to generate dialogue with unique personalities, strategize gambits based on their roles, reason out what other players (AI or human) are hiding, and then vote for the most suspicious person (or the werewolf’s scapegoat). 

They then set the Google AI bots loose, testing how good they were at spotting lies or how susceptible they were to gaslighting. They also tested how the LLMs did when removing specific capabilities like memory or deductive reasoning, to see how it affected the results. 

A slide during the GDC 2024 panel "Simulacra and Subterfuge: Building Agentic 'Werewolf'". It shows an example of the Generative Werewolf game in action, with bots attempting to deceive villagers or sniff out werewolves.

(Image credit: Michael Hicks / Android Central)

The Google engineering team was frank about the experiment’s successes and shortcomings. In ideal situations, the villagers came to the right conclusion nine times out of 10; without proper reasoning and memory, the results fell to three out of 10. The bots were too cagey to reveal useful information and too skeptical of any claims, leading to random dogpiling on unlucky targets.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *