
TL;DR: OpenRelay is building an Inference Delivery Network (IDN) on top of its CDN for GPUs. We provide one entrypoint for inference across every chip.
Send us a workload and we’ll cut your costs up to 20% by running it through our network, balancing the best available accelerator (NVIDIA, TPU, Trainium, AMD) across clouds and handing back a single endpoint. You never pick the hardware, chase quotas, or rebuild your stack per provider. And if you run GPUs or other AI accelerators (an NVIDIA cluster, TPUs, a reserved fleet), you can plug your capacity into our rails and we'll bring the inference demand. Get started here.
P.S. sign up for our Luma page to be notified of our Launch Party!
Hey guys! We're Jaden and Prashant, the founders of OpenRelay. We came at the same problem from opposite sides.
We kept seeing the same thing: there's more AI compute than ever, and you still can't reach it. We left our jobs to fix that.
Compute is everywhere; liquidity is nowhere. Accelerators are scattered across dozens of clouds, chip vendors, and operators, each behind its own quotas, drivers, contracts, and console. Teams that want to run inference can't reach the capacity, and the capacity can't reach them.
Running inference naively is easy; running it at scale is hard. Any operator can spin up a model on a box. Turning idle GPUs into a production endpoint (load balancing, isolation, autoscaling, failover, metering, billing) is a software platform most individual operators won't build. So capacity sits stranded behind their own front door, and developers either overpay for scarce reserved GPUs or hand-stitch a fragile multi-provider stack.
We built one set of rails across all of it.
Send a workload (any container or model) and we schedule it onto the best-fit accelerator, attach a production endpoint, and scale it. We route to the cheapest accelerator that meets your latency and throughput targets, and we benchmark continuously. If the market can't beat renting yourself, we backstop with our own inference demand, so you’re never on the hook.
You never pick the chip, cloud, or region; we handle routing, isolation, failover, metering, and billing across 4+ clouds and every major chip family. We're language- and framework-agnostic (no SDK lock-in), and you can drive everything from our CLI and REST API.
It's two-sided: if you run GPUs or other accelerators (a neocloud, a data center, a pay for an underutilized reserved fleet), you can connect that capacity to our rails and we'll bring the inference demand, turning hardware into a revenue-generating endpoint without building a platform yourself.
We're live in production, generating 100 billion tokens a week across 22 physical locations across Europe, APAC, North America, and the Middle East running on 8 different accelerator SKUs.
We'd love to hear from you!