TL;DR
Baud is building a new chip for training and inference of large AI models
We have developed a new arithmetic representation of neural networks that does not contain multiplications. Our silicon is architected specifically for this representation and doing this lets us train and serve AI faster and more efficiently than incumbents.
Our Team
The Problem
We are living in exciting times where humanity has figured out a way to turn sand and sunlight into machines of scientific discovery.
But creating and serving these machines aka frontier AI models, is extremely capital intensive.
NVIDIA’s publicly documented training recipe of the Nemotron model took 6144 H100 GPUs, going on for around 3 months, drawing around 4.3 megawatts of electricity.
That's $44M in rentals or $245M in capex plus an ungodly number of developer hours standing up distributed training. Models like Claude Opus 4.6 surprised the world, and that's just the beginning. But at these numbers, only a handful of teams on Earth control the frontier in AI and have a shot at a new opus-level moment in other domains like robotics, world models, drug discovery, material science, physics, simulation, the list goes on.
Current attempts at solving this such as Cerebras, Groq, Etched, D-Matrix and 50 other chip startups are only attacking the silicon side and leaving the fundamental math side of the problem untouched, because they are focused on inference and hence have to run existing model weights.
When we looked deeper, we found that if you let go of the constraint of running existing model weights, a radically different silicon architecture becomes possible, with an order of magnitude better performance.
We think we're just scratching the surface of what's possible when you co-design the hardware with the model architecture.
And that’s what we are betting the farm on.
Solution
We have developed a new way to represent neural network architectures that eliminates multiplications from both the forward and backward passes, and also compresses the weights by more than 10x without any loss in intelligence. The catch is that you need to train the model in this representation, or start with a base model that is also in this representation.
Both during training and inference, our chip does not need multiplier circuits.
This makes our individual cores drastically smaller and simpler than say GPU’s tensor cores or the PEs in TPU’s systolic arrays.
Which means we can pack way more compute and SRAM on the die and move way more weights for the same memory bandwidth.
This makes our chip faster, less power hungry and simpler to build - letting us use cheap off the shelf server hardware, commodity DRAM, standard chip to chip interconnects and air cooling equipment.
We provide a developer platform based on our hardware that makes creation of frontier AI frictionless and economical to serve, so businesses of all sizes can start owning their intelligence.
We are live already
Our first chip is validated on the Global Foundries 12nm process and on schedule for tape-out by the end of this year.
Our training and inference service is already live on a cluster of FPGAs running the chip architecture in emulation.
Using FPGA emulation, we've developed a compiler and a distributed training stack that converts any PyTorch exportable model to our format with bit exact results in most cases, including popular opensource architectures such as GLM, Qwen, Gemma, DeepSeek, Flux, Wan and others.
You can checkout a small demo of our chip running in emulation on a single FPGA here: https://baudlabs.ai/demo
Our partner program
Baud is now working with design partners who can test our systems and help smooth out the developer experience, in exchange for reserved capacity on our first cluster for training and inference.
Capacity is limited. Contact us if you are training something new or are tired of paying through the nose for intelligence that you do not own.
https://baudlabs.ai/partner