HomeCompaniesRobocurve

Evals for robots

Robocurve builds open-source tools and independent benchmarks to measure how well robots can do real-world jobs. Instead of relying on unverified demo videos from frontier labs, we score their models on reproducible benchmarks that anyone can trust. Today there are no well-run, standardized robotics benchmarks. Labs evaluate in-house, and no independent group has stepped in to run continuous benchmarking as a service. The result is that no one actually knows how good anyone else is, or where the real frontier sits. Good benchmarks require operating and maintaining physical hardware and real-world setups. Simulation only goes so far, since models that look strong in sim can show large performance gaps once deployed in the real world. Building real-world benchmarks means coordinating job-domain experts, evals engineering, and hands-on robotics all at once. Frontier labs are targeting general-purpose robotics by 2028, yet the field of robotics evals barely exists. Whoever builds the trusted measure of robot capability becomes the reference everyone relies on. We combine backgrounds in AI evals and robotics to build and grow this field as fast as possible. We already shipped v1 of our open-source framework, Inspect Robots, and ran our first pilot scoring a frontier model on a real robot.
Active Founders
Jay Chooi
Jay Chooi
Founder/CEO
Founder and CEO of Robocurve. MA Statistics and BA CS/Math from Harvard. Jay builds evals of frontier AI and forecasts its progress. Previously Research Fellow at MATS, researcher at the UK AI Security Institute, and top contributor to Inspect Evals, the UK government's AI eval framework. Published at ACM EC, ICML, ACL, and EMNLP. Called the 2024 election correctly in all 50 states. Won a Rhodes Scholarship. Gold medal at the International Olympiad on Astronomy and Astrophysics.
Company Launches
Robocurve — Real-World Evaluations of Physical AI
See original launch post

Hi, I’m Jay, co-founder of Robocurve. We measure and reports what frontier robots can actually do in the real world.

Our launch video https://youtu.be/Cjkt2ikvsiQ

Frontier labs are racing to build general-purpose robots within two years, yet no one knows what today's robots are truly capable of. Demo videos of robots performing narrow, cherry-picked tasks proliferate, but there are no rigorous, continuous, standardized benchmarks for robotics.

Robocurve builds benchmarks for robots and the AI models that control them, to help the world understand how far robotics has come and how fast it's moving. Our benchmarks range from playing Jenga and making a sandwich to constructing a data center.

General-purpose robots could arrive before the end of the decade, with far-reaching implications: accelerated data center buildouts, explosive economic growth from a fully automated economy, and intensified geopolitical competition between the US and China. Real-world evaluations from Robocurve will help clarify the timeline to general-purpose robots so that society can prepare for their arrival well ahead of time.

Jay is a Rhodes Scholar and a top contributor to UK AI Security Institute's Inspect Evals. He has published papers in ICML, EMNLP, ACL, ACM EC and his research has been featured in MIT Technology Review. Previously, he spent time at AWS Bedrock, MATS, and Harvard where he completed his BA and MA with a highest honors thesis on AI and democracy.

If you build robots or AI models, reach out! We'd love to put them to the test. Learn more at robocurve.org.

Robocurve
Founded:2026
Batch:Summer 2026
Team Size:3
Status:
Active
Location:San Francisco
Primary Partner:Ankit Gupta