Google DeepMind has added two new games to its AI model benchmarking platform, "Kaggle Game Arena": Werewolf and Poker.

While the platform previously only featured Chess, it can now measure decision-making under imperfect information. This is because real-world judgment differs from the perfect information available on a chessboard.

Chess has been used to evaluate strategic reasoning and long-term planning. In contrast, Werewolf is a team-based game that progresses solely through natural language dialogue. Models must distinguish between lies and truth to identify the hidden "Werewolf."

This evaluation measures "soft skills," such as negotiation and the ability to handle ambiguity, which are required for AI assistants. It also serves as a controlled environment for research into agent safety.

Currently, Gemini 3 Pro and Gemini 3 Flash hold the highest Elo ratings on the Chess leaderboard. The performance improvement over the previous generation, Gemini 2.5, highlights the rapid pace of model advancement.


Source: Advancing AI Benchmarking with Game Arena (HN 134pt, 54 comments) (HN Search (backfill), 2026-02-03)