Today, we're introducing Idler – a frontier data research lab. We build the evals and environments that the world's leading frontier labs use to measure and train their models.
In the last year, we’ve developed thousands of coding evals and environments for a top lab, continuously improving our product to keep pace with the frontier – even as models have gotten 8x better at long-horizon coding.
Now, we're offering off-the-shelf datasets across an even broader set of capabilities: long-horizon software engineering, cybersecurity, recursive self-improvement, legal, long-horizon strategy & ops, enterprise safety, and many others.
ShelfLife, our first public benchmark, is out today — a digital twin seeded with the complete operational data of a live, profitable, growing multi-brand retailer running 9 stores, 3 commerce platforms, and a multi-supplier sourcing operation. Post-training Nemotron 3 Nano on 41 ShelfLife tasks raised its Finance Agent Benchmark score from 30.0% to 41.8% (p ≤ 0.001); full results and sample tasks are available upon request.
Our ask. Reach out if:
We’d love to hear from you at hello@idler.ai. Thanks for the support!