
TL;DR: As many researchers and labs will say, there is a massive shortage of genuinely good AI evaluations. Multi-agent environments are an underdeveloped way of evaluating models that we think will help address our increasingly worse shortage of evals. We're starting with Multi-Agent Arena, where anyone can come compete against frontier AI models in social multiplayer games, as we build towards the full stack for multi-agent environments. Try Multi-Agent Arena for free at https://olamlabs.ai/ :)
https://drive.google.com/file/d/186GgmnOLz4Yc1BpFk6AxHd3aqUdFYCQy
Multi-agent environments provide an essential, underutilized framework for benchmarking autonomous AI models. Multi-Agent Arena is an interactive platform where human participants compete against frontier AI models in multiplayer games, serving as an experimental foundation to evaluate complex agent interactions. Multi-Agent Arena is the first of many stepping stones to building the entire stack and ecosystem for multi-agent environments.
The Evaluation Gap
Current AI evaluations primarily focus on single-agent, task-oriented capabilities. While benchmarks such as SWE-bench measure functional task execution, they do not assess the dynamic social behaviors required for autonomous real-world deployment: negotiation, coalition stability, long-horizon planning against adaptive opponents, and spontaneous deception.
Simulated multi-agent environments induce these complex interactions while preserving complete experimental control. Unlike unstructured real-world deployment, these environments offer full data legibility, enabling researchers to inspect every reasoning trace, state transition, and stochastic variable behind model decisions.
Our Preliminary Evaluations
The platform runs deterministic, replayable environments hosting human players alongside anonymized frontier models. Structured as accessible multiplayer games, the infrastructure captures interaction data to score models across two primary axes in our first evaluated game, Poker:
Initial Metric: Deception Index
The benchmark's initial diagnostic metric, the Deception Index, quantifies deliberate misrepresentation per 10,000 turns. A deceptive event is logged only when an agent's private chain-of-thought confirms awareness of a true state while its external communication directly contradicts it. In baseline tests with no explicit instructions or prompts to deceive, frontier models demonstrated performance variance spanning nearly two orders of magnitude across identical environments.
Complete evaluations, research methodology and technical documentation are available at olamlabs.ai/research/social-arena
Launch
We already passed a thousand matches a day in our early access, are nearing a billion tokens used per day, ranking on the OpenRouter leaderboards, and have thousands of users that are paying to play in the arena while we work with researchers on use cases they have for our work and environments. We’re hoping this is just a stepping stone and want to scale towards our roadmap and ambitions of what we believe multi-agent environments, when properly developed with the right stakeholders in the training/eval ecosystem, can achieve for AI progress.
What’s in it for you?
Om Buddhdev (CEO) — Previously a Product Engineer then promoted to Staff Engineer at AI Dungeon / Latitude working on AI infra, agents, and consumer platform work, left for Olam Labs. Before that he was among the best Valorant players in the world (top 25, 8x radiant), at 17 a large content creator on TikTok with 195k followers, and as a young teenager made Roblox games played millions of times, among other things in his past. More on Om at https://www.sensho.xyz/
Shreshth Sharma (CTO) — Background in production data engineering: was a Data Scientist at RBC while doing a double degree at the University of Waterloo in CS and Math, dropped out of both for Olam Labs. He has a background in many data and statistics projects, alongside briefly being a top 50 Fifa player. More on Shreshth at https://shresh.ca/
Om and Shreshth have been friends since middle school, and Olam Labs was born out of their personal interest and desire to build out the entire stack for multi-agent environments to be the default way we evaluate models in the future.
Om and Shreshth met in middle school. They were the only two kids in the “gifted” program in their grade, and their first project together was running a snacks and candy store at their school where the profits were used to donate to charity.
During Christmas holiday in 2025, Om created a multi-agent simulated environment where agents ran their own company in markets competing against each other, and showed it to Shreshth.
Shreshth, with a background in data and RL among other things, loved the implications of it. Om discussed with him how if multi-agent environments were developed and productionized at scale via frameworks, APIs for infra, and interactive games the kind of data and evaluations it could generate.
The problem was (and still is) that despite the success of multi-agent environments in research, the ecosystem around them has not been pursued to a degree where the utility of multi-agent environments is made readily available to everyone.
They decided to pursue it and applied to YC, and have been in pursuit of a future where we can understand and measure the models that we work with everyday and that will become increasingly autonomous in our everyday lives with AI acceleration. Om and Shreshth both believe there is a severe shortage of good AI evaluations, especially on qualitative behaviors, of which multi-agent environments (and allowing others to build them) is very well suited to solve.
More on Shreshth: shresh.ca Shresh’s X: https://x.com/shresh
More on Om: sensho.xyz Om’s X: https://x.com/sensho