
The Idea
Advancement in AI relies on chips. Transformers won because they mapped well to GPUs. Frontier Labs are now co-designing their own silicon to maximize intelligence/joule.
It takes years and 100s of experts to bring a chip from architecture through tape-out.
Our goal as a company is to push the frontier of AI on chip design tasks - enabling advanced chips to be taped-out by small teams in months.
We also believe this is the path forward to achieving true recursive self-improvement (RSI). Where AI designs better chips and better chips train more intelligent AI.
To achieve this, we're building RL Environments For Long-Horizon Chip Design.
Our environments enable 1000s of parallel rollouts for training on chip design tasks. We also curate a private held-out test set for performing evals in the environment to compare model capabilities.
What Exists
Chip design should be a perfect domain for this type of RL. Instead:
- The academic benchmarks are one-shot puzzles of ~100 lines, and they're in every training set.
- The field's de facto standard (NVIDIA's CVDP) has saturated - going from 34% to 97% once agents could iterate against the grader.
- The published tasks are short-horizon and graded by running testbenches. In one study, designs that passed their bundled testbenches 95-97% of the time were only ~20% correct under formal checking.
And none of it resembles the real loop: months of work, tools in the loop, and hard trade-offs between latency, throughput, area, and power.
What We Build
Our RL Environment is designed for long-horizon chip design tasks.
- Long-horizon. One episode is the real loop: choose a microarchitecture, write the RTL, simulate, synthesize, check timing and power, revise.
- Reliable grading. Every task is an executable contract with exact checking. No LLM judge anywhere in the reward path. Correctness is a hard gate: a fast, small, wrong chip earns zero.
- Physics as the yardstick. Every task ships with a computed speed-of-light bound: the best the physics of the task allows. Scores are reported as a fraction of that bound, so they're comparable across tasks.
- The full trade-off. Real chips are trade-offs. We record area, power, latency, and throughput as an unreduced vector - labs can scalarize this into a terminal reward for training however they want.
- Deterministic. Every graded result replays bit-identically on another machine.
Why Now
Chip Design is increasingly something the frontier labs care about. They are all designing their own custom inference chips. Anthropic is already hiring chip-design RL engineers at $500-850k to build these types of environments in-house - and there’s currently no yardstick that hasn’t been saturated.
Who We Are
I was the first intern & youngest hire into GPU Architecture at Apple, working with industry-leading architects.
I've seen how leading-edge chips get designed end-to-end: what the loop looks like, and how the power/perf/area trade-offs actually get graded.
Our environments run that same loop & formalize that grading.
Asks
If you’re working on RL, post-training, or data at a lab or neolab, or interested in post-training models for chip-design capabilities. Please reach out, founders@standardmachines.com.
Read our full launch post, here.
- Jacob