TL;DR: We’re automating robot evaluation– the loop where you wait for your robot to roll out, judge its success, and then reset the scene. Today, we’re launching step one: a success detector that automatically tells you whether your robot actually completed its task. Book a demo with founders@instancelabs.ai and we’ll visit your office to show you how it works!
The problem
In Lucy’s robotics research this past year, she spent way too much time being a babysitter for her robot. Her project was to get a humanoid to find objects around a room, so her days went like this: start a rollout, watch the robot open all the drawers in the room for 30 minutes, mark success/fail in a spreadsheet, then get up, close all the drawers, and put the objects back where they started. Then do it again, and again, and again. Some days ended with back pain, and for weeks she couldn’t work on anything else because she was stuck doing exactly this.
That was one robot, in one lab, running a single policy. Industry teams run fleets of robots across dozens of policies, with multiple people assigned to sit and watch each one. Every robotics team we’ve talked to has some version of this– and it will only grow as robot foundation models scale.
Our solution
Instance is building the autonomous evaluation rig for robot learning: your robot rolls out, we judge success, a second robot resets the scene, and the next rollout starts– with no humans in the loop at all.
Today, we’re starting with the success judge. Point us at a task description and your video rollouts, and we return a verdict with detailed subtask captions. We benchmarked our verifier against 10,000+ held-out, human-labeled episodes across 7 robot platforms, and it resulted in higher accuracy than Claude Opus 4.8, at a fraction of the latency.
Try it
Try it at demo.instancelabs.ai – drop in a video, or record one from your phone and watch it get verified live.
Team
We’re Claire and Lucy, MIT grads and friends since middle school!
Claire studied math and CS at MIT. At NASA JPL, she built software to simulate planetary atmospheres, and at the MIT Media Lab, she researched sustainable propulsion for spaceflight. She loves working where software meets the physical world.
Lucy studied CS at MIT and completed her MEng in MIT's Learning and Intelligent Systems lab. She’s an experienced robotics researcher, and has in the past worked on satellite software at SpaceX, LLM unit testing at Amazon, and brain computer interfaces at Blackrock Neurotech.
Our ask
If you are:
Email us (founders@instancelabs.ai) to book a demo. We’d love to swing by your office with sweet treats to show you how it works!