For people interested in these kinds of benchmarks, I have two multiplayer, multi-round games: - E...

zone411 • today at 1:38 AM • 0 replies • view on HN

For people interested in these kinds of benchmarks, I have two multiplayer, multi-round games:

- Elimination Game Benchmark: Social Reasoning, Strategy, and Deception in Multi-Agent LLM Dynamics at https://github.com/lechmazur/elimination_game/

- Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure at https://github.com/lechmazur/step_game/

alt Hacker News