HomeCompaniesInstance

Automated evals for robot policies

We’re automating robot evaluation– the loop where you wait for your robot to roll out, judge its success, and then reset the scene. We're starting with step one: a success detector that automatically tells you whether your robot actually completed its task along with detailed subtask breakdown. We're MIT computer scientists & friends since middle school, with backgrounds at NASA JPL, SpaceX, and AWS. Book a demo with founders@instancelabs.ai.
Active Founders
Claire Mao
Claire Mao
Co-founder and CEO
MIT Math + CS. Prev @ NASA (JPL), BCG, MIT Media Lab
Lucy Cai
Lucy Cai
Cofounder and CTO
MIT CS + AI. Robotics research @ MIT LIS Lab; prev @ SpaceX, Amazon (AWS)
Company Launches
Instance: Automated evaluation for robot policies, starting with the success detector
See original launch post

TL;DR: We’re automating robot evaluation– the loop where you wait for your robot to roll out, judge its success, and then reset the scene. Today, we’re launching step one: a success detector that automatically tells you whether your robot actually completed its task. Book a demo with founders@instancelabs.ai and we’ll visit your office to show you how it works!

https://youtu.be/EqxkPffnUKk

The problem

In Lucy’s robotics research this past year, she spent way too much time being a babysitter for her robot. Her project was to get a humanoid to find objects around a room, so her days went like this: start a rollout, watch the robot open all the drawers in the room for 30 minutes, mark success/fail in a spreadsheet, then get up, close all the drawers, and put the objects back where they started. Then do it again, and again, and again. Some days ended with back pain, and for weeks she couldn’t work on anything else because she was stuck doing exactly this.

That was one robot, in one lab, running a single policy. Industry teams run fleets of robots across dozens of policies, with multiple people assigned to sit and watch each one. Every robotics team we’ve talked to has some version of this– and it will only grow as robot foundation models scale.

Our solution

Instance is building the autonomous evaluation rig for robot learning: your robot rolls out, we judge success, a second robot resets the scene, and the next rollout starts– with no humans in the loop at all.

Today, we’re starting with the success judge. Point us at a task description and your video rollouts, and we return a verdict with detailed subtask captions. We benchmarked our verifier against 10,000+ held-out, human-labeled episodes across 7 robot platforms, and it resulted in higher accuracy than Claude Opus 4.8, at a fraction of the latency.

Try it

Try it at demo.instancelabs.ai – drop in a video, or record one from your phone and watch it get verified live.

Team

We’re Claire and Lucy, MIT grads and friends since middle school! 

Claire studied math and CS at MIT. At NASA JPL, she built software to simulate planetary atmospheres, and at the MIT Media Lab, she researched sustainable propulsion for spaceflight. She loves working where software meets the physical world.

Lucy studied CS at MIT and completed her MEng in MIT's Learning and Intelligent Systems lab. She’s an experienced robotics researcher, and has in the past worked on satellite software at SpaceX, LLM unit testing at Amazon, and brain computer interfaces at Blackrock Neurotech.

Our ask

If you are:

  • training robot policies
  • running evals
  • deploying robots in the real world
  • or can intro us to someone who is,

Email us (founders@instancelabs.ai) to book a demo. We’d love to swing by your office with sweet treats to show you how it works!

uploaded image

YC Photos
Instance
Founded:2026
Batch:Summer 2026
Team Size:2
Status:
Active
Location:San Francisco
Primary Partner:Ankit Gupta