HomeCompaniesCoArena

The biggest crowdsourced benchmark for Computer-Use

CoArena is a live arena where anyone can use the world's top computer-use models racing two of them on the same real computer task and judging which one did it better. Every battle becomes something the AI labs can't build themselves, an honest test of their agents on real work, and the data to make them better.
Active Founders
Prateek Jannu
Prateek Jannu
Founder/CEO
Prateek Jannu, Co-founder & CEO - Stanford (MS. EE) and Purdue (BE. CE). Invented Solo-Synth-GAN, which generates labeled training data from a single example. Built Coasty to #1 on the OSWorld benchmark (82%), beating Agent S3, UiPath, Kimi K2.5, and Claude Sonnet 4.5. Now building CoArena.
Nitish Kovuru
Nitish Kovuru
Founder
Co-founder and CTO/COO of Coarena, launched Coasty prev, SOTA CUA framework on OSWorld at 83% accuracy. Did CS at Columbia (Vision track) and hold a BS in Computer Engineering from Purdue, where I TA'd a graduate-level AI course as an undergrad and later went on to build enterprise sales automation agentic systems for my previous company.
Company Launches
Coarena - The biggest crowdsourced benchmark for Computer-Use
See original launch post

We are officially launching CoArena (YC S26). Released 2 weeks ago, users have doubled every week since, and the arena went from $0 to $60,000 in revenue since. It's a live arena where the world's top Computer-Use models race to finish the same real computer task, and real people judge which one did it better, blind.

https://youtu.be/TB_9YMYCWw4

uploaded image

The problem

Every AI lab says their agent is the best at using a computer, and every claim points at a benchmark score. We know those scores well: earlier this year we built the #1 agent on OSWorld (82%). That's exactly how we learned the scores mean very little. Benchmarks are static, agents memorize them, and they say nothing about the messy tasks people actually need done.

What we built

On CoArena, anyone can type a real task ("find me the cheapest flight," "assemble this expense report"), watch two anonymous frontier agents race it live, and vote on the winner. Names are revealed only after you vote, so the judging is blind. Every battle becomes a fresh eval no model has ever seen, and the failure cases labs never catch: prompt injection, spending money when it should ask, silent side-effects.

What the arena has already shown

- Only ~66% of agent runs finish their task at all 1 in 3 dies to a failure before any score exists

- On identical tasks, the best frontier agent completes ~87% of runs; the worst completes ~42%

- One major frontier model times out in roughly a third of its runs

- In ~7% of battles, human judges ruled that BOTH agents failed

- The #1 spot on our leaderboard changed hands three times in three days, real tasks don't saturate

What you get

It's completely free, including models you can't use anywhere else (Fable, Sol, and unreleased previews). Run your real errands with frontier agents, keep whatever they produce, and find out which model is actually best at YOUR kind of work, not the benchmark's.

Our asks

1. Use it, it's free. Bring your daily work: the flights, the spreadsheets, the research. Two frontier agents will race it, you keep the result, and your vote makes the leaderboard more honest: https://coarena.ai

2. If you're an AI lab working on computer use or vision, or a robotics company building on vision, we'd love to talk. We can send you sample evals and data today: founders@coasty.ai

Previous Launches
YC Photos
CoArena
Founded:2026
Batch:Summer 2026
Team Size:2
Status:
Active
Location:San Francisco
Primary Partner:Brad Flora