{"id":109432,"title":"Can you beat frontier LLMs at social strategy games? | Multi-Agent Arena by Olam Labs, evaluating through multi-agent simulations","tagline":"Play games like Risk, Poker, or Codenames against frontier AIs. We're running public long-horizon games to do evaluations on social behavior and agentic performance.","body":"**TL;DR:** As many researchers and labs will say, there is a massive shortage of genuinely good AI evaluations. Multi-agent environments are an underdeveloped way of evaluating models that we think will help address our increasingly worse shortage of evals. We're starting with **Multi-Agent Arena**, where anyone can come compete against frontier AI models in social multiplayer games, as we build towards the full stack for multi-agent environments. Try Multi-Agent Arena for free at \u003chttps://olamlabs.ai/\u003e :) \\\n\\\n[https://drive.google.com/file/d/186GgmnOLz4Yc1BpFk6AxHd3aqUdFYCQy](https://drive.google.com/file/d/186GgmnOLz4Yc1BpFk6AxHd3aqUdFYCQy/view?usp=sharing)\n\n---\n\n## **Overview**\n\nMulti-agent environments provide an essential, underutilized framework for benchmarking autonomous AI models. Multi-Agent Arena is an interactive platform where human participants compete against frontier AI models in multiplayer games, serving as an experimental foundation to evaluate complex agent interactions. Multi-Agent Arena is the first of many stepping stones to building the entire stack and ecosystem for multi-agent environments.\\\n\\\n**The Evaluation Gap** \\\nCurrent AI evaluations primarily focus on single-agent, task-oriented capabilities. While benchmarks such as SWE-bench measure functional task execution, they do not assess the dynamic social behaviors required for autonomous real-world deployment: negotiation, coalition stability, long-horizon planning against adaptive opponents, and spontaneous deception.\n\nSimulated multi-agent environments induce these complex interactions while preserving complete experimental control. Unlike unstructured real-world deployment, these environments offer full data legibility, enabling researchers to inspect every reasoning trace, state transition, and stochastic variable behind model decisions.\n\n**Our Preliminary Evaluations**\\\nThe platform runs deterministic, replayable environments hosting human players alongside anonymized frontier models. Structured as accessible multiplayer games, the infrastructure captures interaction data to score models across two primary axes in our first evaluated game, Poker:\n\n* **Competitive Capability:** One of our preliminary methods is using a Bradley-Terry model to yield Elo ratings with bootstrapped confidence intervals, supplemented by domain-specific metrics such as big blinds per 100 hands in poker. Top-quartile human participants currently outperform all benchmarked models in Poker.\n* **Behavioral Diagnostics:** Model reasoning traces are evaluated against ground-truth game states using expert-curated rubrics and variance-checked grader panels to ensure inter-annotator agreement.\n\n**Initial Metric: Deception Index** \\\nThe benchmark's initial diagnostic metric, the Deception Index, quantifies deliberate misrepresentation per 10,000 turns. A deceptive event is logged only when an agent's private chain-of-thought confirms awareness of a true state while its external communication directly contradicts it. In baseline tests with no explicit instructions or prompts to deceive, frontier models demonstrated performance variance spanning nearly two orders of magnitude across identical environments.\n\nComplete evaluations, research methodology and technical documentation are available at [olamlabs.ai/research/social-arena](https://olamlabs.ai/evaluations)\\\n\\\n**Launch**\\\n\\\nWe already passed a thousand matches a day in our early access, are nearing a billion tokens used per day, ranking on the OpenRouter leaderboards, and have thousands of users that are paying to play in the arena while we work with researchers on use cases they have for our work and environments. We’re hoping this is just a stepping stone and want to scale towards our roadmap and ambitions of what we believe multi-agent environments, when properly developed with the right stakeholders in the training/eval ecosystem, can achieve for AI progress.\\\n\\\n\\\n**What’s in it for you?**\n\n1. **Play.** It's free, and every match directly informs the evals. Bring friends, try to beat the models, it’s harder than you'd think. 🃏\n2. **Speak with researchers.** We love speaking with researchers interested in this space and it’s all that we think about; if that’s you please feel free to reach out to us at [founders@olamlabs.ai](mailto:founders@olamlabs.ai) or on \u003chttps://olamlabs.ai/about\u003e \\\n   \\\n   🫡\n\n---\n\n## About the team\n\n**Om Buddhdev (CEO)** — Previously a Product Engineer then promoted to Staff Engineer at AI Dungeon / Latitude working on AI infra, agents, and consumer platform work, left for Olam Labs. Before that he was among the best Valorant players in the world (top 25, 8x radiant), at 17 a large content creator on TikTok with 195k followers, and as a young teenager made Roblox games played millions of times, among other things in his past. More on Om at \u003chttps://www.sensho.xyz/\u003e\n\n**Shreshth Sharma (CTO)** — Background in production data engineering: was a Data Scientist at RBC while doing a double degree at the University of Waterloo in CS and Math, dropped out of both for Olam Labs. He has a background in many data and statistics projects, alongside briefly being a top 50 Fifa player. More on Shreshth at \u003chttps://shresh.ca/\u003e\\\n\\\nOm and Shreshth have been friends since middle school, and Olam Labs was born out of their personal interest and desire to build out the entire stack for multi-agent environments to be the default way we evaluate models in the future.","slug":"ST2-can-you-beat-frontier-llms-at-social-strategy-games-multi-agent-arena-by-olam-labs-evaluating-through-multi-agent-simulations","created_at":"2026-08-05T21:32:50.566Z","updated_at":"2026-09-19T09:41:26.251Z","total_vote_count":9,"url":"https://www.ycombinator.com/launches/ST2-can-you-beat-frontier-llms-at-social-strategy-games-multi-agent-arena-by-olam-labs-evaluating-through-multi-agent-simulations","share_image_url":"//bookface-static.ycombinator.com/assets/ycdc/yc-og-image-c440a0ad1dacfb86eeeb343717479cc54d256614449b4ef719977a0a451f8bc8.png","company":{"id":32130,"name":"Olam Labs","slug":"olam-labs","url":"https://olamlabs.ai","logo":"https://bookface-images.s3.amazonaws.com/small_logos/d6d9543a0527b0b5ac428e043f211c75ac5d1d59.png","batch":"Summer 2026","industry":"B2B","tags":["Artificial Intelligence","Reinforcement Learning","Gaming","Data Engineering"],"search_path":"https://bookface.ycombinator.com/company/32130"}}