HomeCompaniesLamb Labs

Lightning fast chips with hardcoded AI models

Lamb Labs is building MPUs (Model Processing Units), custom chips that hardcode LLM weights directly into the chip. Unlike GPUs, which move weights between memory and compute, MPUs keep the model on-chip, targeting up to 20,000+ tok/s and a 63× higher intelligence per watt.
Active Founders
Niki Kotecha
Niki Kotecha
Founder/CEO
Founder at Lamb Labs. AI PhD at Imperial College London + I-X Scholar. MEng/BA in Engineering from University of Cambridge. Alan Turing Institute Global Fellowship and 160+ citations
Thomas Lanning
Thomas Lanning
Founder/CTO
Founder at Lamb Labs. Masters in Mathematical and Theoretical Physics from Oxford. Top of his class in undergraduate at Edinburgh University. (Re)discovered the Higgs-boson using ML and CERN data.
Company Launches
Lamb Labs: Custom Chips for AI Inference
See original launch post

Hi Everyone! 

We’re Niki and Thomas - cofounders of Lamb Labs.  

TL;DR: We build custom AI inference chips that run inference at 20,000+ tok/s and 63x higher intelligence per watt than a traditional GPU. Fully local and private by construction, no subscriptions, no API costs. 

Our mission: the world’s fastest LLM at the lowest power. 

https://youtu.be/MhhY1Dy9ltE?feature=shared

Our hot take: AI models are finally getting good enough to hardcode them directly into silicon. Once you stop treating the chip as a general-purpose GPU and burn a specific model into hardware, inference gets radically faster and cheaper per watt. Most of the industry is still building flexible, general chips. We think that flexibility is exactly what's wasting your energy. So we’re on a mission to reduce the energy footprint of AI. 

The Problem

AI is taking power on two fronts: concentrating control of data and intelligence, and consuming ever more electricity to grow. That future scares us, so we're building against it.

Here's the technical root of the waste: GPUs burn most of their energy during inference moving data, not computing. LLM/VLM inference is memory-bound, fetching a weight from memory costs ~100–1000x the energy of the arithmetic you do with it. So you underutilize the GPU you paid for: some GPUs only hit 20–40% compute utilization during LLM inference while the rest idles waiting on memory.

Cloud hosting adds its own tax: expensive per-token bills, latency you can't control, network dependencies, and your data leaving your control.

Our Solution

We build the model and the chip together. 

The Hardware Layer

  • Custom architecture, RL-tuned. We built an RL environment that co-designs the chip architecture for a given model, optimizing directly for speed and energy per token.
  • Aggressive quantization. Removes complex math, which shrinks the size and cost of the chip.
  • Weights on-chip. The model weights live where the compute is overcoming the memory-bandwidth bottleneck

 

Because everything is hardcoded, it's fully local and secure by construction. No connectivity dependency, nothing leaves the device.

The Software Layer 

Designing the chip led us to solving a fundamental flaw in LLMs so we are now open-sourcing a new model next week, converted from another open source model, to run 2x faster on any hardware. Come back next week for part 2 of our launch to find out more….

Our Ask

If you answer "yes" to any of these, we'd love your help:

  • Are you paying for inference in watts, not just dollars? Neocloud or private DC operators who are power- or hardware-cost-bound rather than demand-bound, talk to us.
  • Do you need inference where the cloud isn't an option? Private AI deployments, regulated, air-gapped, sovereign, or bandwidth-starved deployments.
  • Do you want to deploy the fastest AI models in your stack? High frequency trading, agentic workflows 
  • Are you stuck on a thermal or battery budget? Robotics, wearables, consumer AI hardware where a Jetson-class part is over-provisioned and runs too hot.
  • Can you introduce us to anyone above? Warm intros are the single most useful thing.

Contact us at contact@lamb-labs.com

YC Photos
Lamb Labs
Founded:2026
Batch:Summer 2026
Status:
Active
Location:San Francisco
Primary Partner:Tyler Bosmeny