HomeCompaniesOlam Labs
Olam Labs

Building multi-agent simulations for model evals and training.

Multi-agent environments let us evaluate and train models in complex, simulated worlds. We work with researchers on both evaluations and training for character and agentic performance typically difficult to assess for using current popular datasets. Play one of our first releases, Multi-Agent Arena (https://olamlabs.ai/arena), where humans come play social strategy games against multiple AI agents. It's currently top 50 on OpenRouter and has thousands of matches played each day. We use Multi-Agent Arena to build datasets on agentic performance and real-world socialization.
Active Founders
Om Buddhdev
Om Buddhdev
Founder/CEO
I am building the entire stack for multi-agent environments so we can do better AI evaluations and training. I'm a tech enthusiast and optimize for following my curiosity. In the past I've done some fun things before Olam Labs ranging from esports, to consumer, to engineering, to research. I was one of the best Valorant players in the world, a Staff Engineer at AI Dungeon, done math research, and other stuff. You can read more here: https://www.sensho.xyz/
shreshth sharma
shreshth sharma
Founder/CTO
currently building mulit-agent envs to create better evals @ olam labs, previously: was doing a cs/math at waterloo and business double degree at laurier (dropped out to start olam labs ) creating reinforcement learning systems and doing data science @ RBC did algorithmic arbitrage on prediction markets briefly top 50 on fifa mobile
Company Launches
Can you beat frontier LLMs at social strategy games? | Multi-Agent Arena by Olam Labs, evaluating through multi-agent simulations
See original launch post

TL;DR: As many researchers and labs will say, there is a massive shortage of genuinely good AI evaluations. Multi-agent environments are an underdeveloped way of evaluating models that we think will help address our increasingly worse shortage of evals. We're starting with Multi-Agent Arena, where anyone can come compete against frontier AI models in social multiplayer games, as we build towards the full stack for multi-agent environments. Try Multi-Agent Arena for free at https://olamlabs.ai/ :)

https://drive.google.com/file/d/186GgmnOLz4Yc1BpFk6AxHd3aqUdFYCQy


Overview

Multi-agent environments provide an essential, underutilized framework for benchmarking autonomous AI models. Multi-Agent Arena is an interactive platform where human participants compete against frontier AI models in multiplayer games, serving as an experimental foundation to evaluate complex agent interactions. Multi-Agent Arena is the first of many stepping stones to building the entire stack and ecosystem for multi-agent environments.

The Evaluation Gap
Current AI evaluations primarily focus on single-agent, task-oriented capabilities. While benchmarks such as SWE-bench measure functional task execution, they do not assess the dynamic social behaviors required for autonomous real-world deployment: negotiation, coalition stability, long-horizon planning against adaptive opponents, and spontaneous deception.

Simulated multi-agent environments induce these complex interactions while preserving complete experimental control. Unlike unstructured real-world deployment, these environments offer full data legibility, enabling researchers to inspect every reasoning trace, state transition, and stochastic variable behind model decisions.

Our Preliminary Evaluations
The platform runs deterministic, replayable environments hosting human players alongside anonymized frontier models. Structured as accessible multiplayer games, the infrastructure captures interaction data to score models across two primary axes in our first evaluated game, Poker:

  • Competitive Capability: One of our preliminary methods is using a Bradley-Terry model to yield Elo ratings with bootstrapped confidence intervals, supplemented by domain-specific metrics such as big blinds per 100 hands in poker. Top-quartile human participants currently outperform all benchmarked models in Poker.
  • Behavioral Diagnostics: Model reasoning traces are evaluated against ground-truth game states using expert-curated rubrics and variance-checked grader panels to ensure inter-annotator agreement.

Initial Metric: Deception Index
The benchmark's initial diagnostic metric, the Deception Index, quantifies deliberate misrepresentation per 10,000 turns. A deceptive event is logged only when an agent's private chain-of-thought confirms awareness of a true state while its external communication directly contradicts it. In baseline tests with no explicit instructions or prompts to deceive, frontier models demonstrated performance variance spanning nearly two orders of magnitude across identical environments.

Complete evaluations, research methodology and technical documentation are available at olamlabs.ai/research/social-arena

Launch

We already passed a thousand matches a day in our early access, are nearing a billion tokens used per day, ranking on the OpenRouter leaderboards, and have thousands of users that are paying to play in the arena while we work with researchers on use cases they have for our work and environments. We’re hoping this is just a stepping stone and want to scale towards our roadmap and ambitions of what we believe multi-agent environments, when properly developed with the right stakeholders in the training/eval ecosystem, can achieve for AI progress.


What’s in it for you?

  1. Play. It's free, and every match directly informs the evals. Bring friends, try to beat the models, it’s harder than you'd think. 🃏
  2. Speak with researchers. We love speaking with researchers interested in this space and it’s all that we think about; if that’s you please feel free to reach out to us at founders@olamlabs.ai or on https://olamlabs.ai/about

    🫡

About the team

Om Buddhdev (CEO) — Previously a Product Engineer then promoted to Staff Engineer at AI Dungeon / Latitude working on AI infra, agents, and consumer platform work, left for Olam Labs. Before that he was among the best Valorant players in the world (top 25, 8x radiant), at 17 a large content creator on TikTok with 195k followers, and as a young teenager made Roblox games played millions of times, among other things in his past. More on Om at https://www.sensho.xyz/

Shreshth Sharma (CTO) — Background in production data engineering: was a Data Scientist at RBC while doing a double degree at the University of Waterloo in CS and Math, dropped out of both for Olam Labs. He has a background in many data and statistics projects, alongside briefly being a top 50 Fifa player. More on Shreshth at https://shresh.ca/

Om and Shreshth have been friends since middle school, and Olam Labs was born out of their personal interest and desire to build out the entire stack for multi-agent environments to be the default way we evaluate models in the future.

YC Photos
Hear from the founders

How did your company get started? (i.e., How did the founders meet? How did you come up with the idea? How did you decide to be a founder?)

Om and Shreshth met in middle school. They were the only two kids in the “gifted” program in their grade, and their first project together was running a snacks and candy store at their school where the profits were used to donate to charity.

During Christmas holiday in 2025, Om created a multi-agent simulated environment where agents ran their own company in markets competing against each other, and showed it to Shreshth.

Shreshth, with a background in data and RL among other things, loved the implications of it. Om discussed with him how if multi-agent environments were developed and productionized at scale via frameworks, APIs for infra, and interactive games the kind of data and evaluations it could generate.

The problem was (and still is) that despite the success of multi-agent environments in research, the ecosystem around them has not been pursued to a degree where the utility of multi-agent environments is made readily available to everyone.

They decided to pursue it and applied to YC, and have been in pursuit of a future where we can understand and measure the models that we work with everyday and that will become increasingly autonomous in our everyday lives with AI acceleration. Om and Shreshth both believe there is a severe shortage of good AI evaluations, especially on qualitative behaviors, of which multi-agent environments (and allowing others to build them) is very well suited to solve.

More on Shreshth: shresh.ca Shresh’s X: https://x.com/shresh
More on Om: sensho.xyz Om’s X: https://x.com/sensho

Olam Labs
Founded:2026
Batch:Summer 2026
Team Size:2
Status:
Active
Location:San Francisco
Primary Partner:Harshita Arora