{"id":111689,"title":"Coarena - The biggest crowdsourced benchmark for Computer-Use","tagline":"We build the most advanced and biggest benchmark for computer use","body":"We are officially launching CoArena (YC S26). Released 2 weeks ago, users have doubled every week since, and the arena went from $0 to $60,000 in revenue since. It's a live arena where the world's top Computer-Use models race to finish the same real computer task, and real people judge which one did it better, blind.\n\n\u003chttps://youtu.be/TB_9YMYCWw4\u003e\n\n![uploaded image](/media/?type=post\u0026id=111689\u0026key=user_uploads/1194042/43ddaa95-1f50-4b30-b0a4-9b085433ec3c)\n\n**The problem**\n\nEvery AI lab says their agent is the best at using a computer, and every claim points at a benchmark score. We know those scores well: earlier this year we built the #1 agent on OSWorld (82%). That's exactly how we learned the scores mean very little. Benchmarks are static, agents memorize them, and they say nothing about the messy tasks people actually need done.\n\n**What we built**\n\nOn CoArena, anyone can type a real task (\"find me the cheapest flight,\" \"assemble this expense report\"), watch two anonymous frontier agents race it live, and vote on the winner. Names are revealed only after you vote, so the judging is blind. Every battle becomes a fresh eval no model has ever seen, and the failure cases labs never catch: prompt injection, spending money when it should ask, silent side-effects.\n\n**What the arena has already shown**\n\n\\- Only \\~66% of agent runs finish their task at all 1 in 3 dies to a failure before any score exists\n\n\\- On identical tasks, the best frontier agent completes \\~87% of runs; the worst completes \\~42%\n\n\\- One major frontier model times out in roughly a third of its runs\n\n\\- In \\~7% of battles, human judges ruled that BOTH agents failed\n\n\\- The #1 spot on our leaderboard changed hands three times in three days, real tasks don't saturate\n\n**What you get**\n\nIt's completely free, including models you can't use anywhere else (Fable, Sol, and unreleased previews). Run your real errands with frontier agents, keep whatever they produce, and find out which model is actually best at YOUR kind of work, not the benchmark's.\n\n**Our asks**\n\n1\\. Use it, it's free. Bring your daily work: the flights, the spreadsheets, the research. Two frontier agents will race it, you keep the result, and your vote makes the leaderboard more honest: \u003chttps://coarena.ai\u003e\n\n2\\. If you're an AI lab working on computer use or vision, or a robotics company building on vision, we'd love to talk. We can send you sample evals and data today: [founders@coasty.ai](mailto:founders@coasty.ai)","slug":"T3R-coarena-the-biggest-crowdsourced-benchmark-for-computer-use","created_at":"2026-08-22T19:00:00.303Z","updated_at":"2026-09-19T03:46:50.055Z","total_vote_count":59,"url":"https://www.ycombinator.com/launches/T3R-coarena-the-biggest-crowdsourced-benchmark-for-computer-use","share_image_url":"//bookface-static.ycombinator.com/assets/ycdc/yc-og-image-c440a0ad1dacfb86eeeb343717479cc54d256614449b4ef719977a0a451f8bc8.png","company":{"id":33738,"name":"CoArena","slug":"coarena","url":"http://coarena.ai","logo":"https://bookface-images.s3.amazonaws.com/small_logos/1543623e6506eaf893b1c43e614688a6b2a32bda.png","batch":"Summer 2026","industry":"B2B","tags":["AIOps","Artificial Intelligence","Machine Learning","Reinforcement Learning","Data Labeling"],"search_path":"https://bookface.ycombinator.com/company/33738"}}