
We're Karan and Thomas, and we're building hiloop: infrastructure for automated research.
TL;DR: we gave two stock coding agents 50 B200s and our infra. They ran 4,188 experiments in two days and beat the published state of the art on Karpathy's autoresearch benchmark. The same machinery found a NanoGPT speedrun candidate that beats the best known result. See full writeup.
Now we want to point it at your hardest problems.
The results
On Karpathy’s autoresearch benchmark, our final recipe reached 0.9016 val_bpb. The previous best published result was Recursive’s 0.9109 on B200 hardware. For NanoGPT, our implementation had 12.54% lower runtime than upstream.
There was no bespoke research agent or elaborate scaffold. We used Claude and GPT-5.5 out of the box. What changed was the system around them.
What hiloop does
Give us a hard task, your current agent or model, and an evaluation criteria. Hiloop runs an autoresearch campaign across training, data, prompts, tools, harnesses, and systems, then returns the best verified improvement. It runs in your cloud or hosted by us.
We’re starting with agent and model training: SFT, post-training, continual learning, and optimization.
https://drive.google.com/file/d/1Rv8HttSHZ-yvSJk6hsWQsqCXWTHJQogK
Who we are
We met at Reducto, where we built its ML, post-training, and platform systems. Karan previously led ML infrastructure at DynamoAI . Thomas was a founding engineer at Crosswise and an MLE at Discord and SoFi.
Two asks
Email founders@hiloop.ai to get in touch with us!