
TL;DR: We're Induction Labs, and today we're announcing research we've been working on: imagination models, a foundation model architecture that unlocks scalable learning from internet video. Our first model, Photon-1, beats a production LLM on internal computer use benchmarks with 30x less training compute, at 3x lower serving cost. [Main Article]
Internet video contains millions of hours of people using computers, doing skilled work, and interacting with the world.
Imagination models can scalably learn from this video to get better at completing tasks and understanding the world.
We tested the architecture with Photon-1, a 106B-A5B MoE transformer pretrained on 18 years of computer screen recordings.
Some Results:
We see imagination models as a path to intelligence that learns by observing the world directly, without a human first translating it into text.
Our Ask
We’re working to scale this method to more kinds of video and looking for exceptional people to work with us! If you’re an exceptional thinker, engineer, or researcher (or know anyone who would be a great fit) - shoot us a message at team@inductionlabs.com