HomeCompaniesOpenRelay
OpenRelay

Distributed, hardware-agnostic AI inference

Active Founders
Jaden Wang
Jaden Wang
Founder/CEO
ex Voltage Park Engineering HPC, ex TensorDock, Founder@Heaviside Compute, UW Dropout
Prashant Patel
Prashant Patel
Founder/CTO
Founder at OpenRelay (YC S26). Previously Staff Engineer at Voltagepark building managed Inference and Orchestration software for 10,000+ GPU clusters and serverless compute that scaled to 10M+ executions/month. Founding member of Amazon Bedrock at AWS, led Custom Model Import from 0 to multi-million ARR and scaled inference infrastructure for models up to trillions of parameters. MS CS from NYU.
Company Launches
OpenRelay: The Inference Delivery Network
See original launch post

TL;DR: OpenRelay is building an Inference Delivery Network (IDN) on top of its CDN for GPUs. We provide one entrypoint for inference across every chip. 

Send us a workload and we’ll cut your costs up to 20% by running it through our network, balancing the best available accelerator (NVIDIA, TPU, Trainium, AMD) across clouds and handing back a single endpoint. You never pick the hardware, chase quotas, or rebuild your stack per provider. And if you run GPUs or other AI accelerators (an NVIDIA cluster, TPUs, a reserved fleet), you can plug your capacity into our rails and we'll bring the inference demand. Get started here.

P.S. sign up for our Luma page to be notified of our Launch Party!

https://luma.com/s6x8cx5x


https://youtu.be/7uNhqcjBXJg

The Team

Hey guys! We're Jaden and Prashant, the founders of OpenRelay. We came at the same problem from opposite sides.

  • Jaden built a half-megawatt data center at 20 from a warehouse shell, then joined Voltage Park early, where he ran virtualization and helped stand up HPC.
  • Prashant built accelerator infra at AWS, where he ran custom model deployments and wrote custom kernels to squeeze out Trainium performance.

 

We kept seeing the same thing: there's more AI compute than ever, and you still can't reach it. We left our jobs to fix that.

The Problem

Compute is everywhere; liquidity is nowhere. Accelerators are scattered across dozens of clouds, chip vendors, and operators, each behind its own quotas, drivers, contracts, and console. Teams that want to run inference can't reach the capacity, and the capacity can't reach them.

Running inference naively is easy; running it at scale is hard. Any operator can spin up a model on a box. Turning idle GPUs into a production endpoint (load balancing, isolation, autoscaling, failover, metering, billing) is a software platform most individual operators won't build. So capacity sits stranded behind their own front door, and developers either overpay for scarce reserved GPUs or hand-stitch a fragile multi-provider stack.

Solution

We built one set of rails across all of it.

Send a workload (any container or model) and we schedule it onto the best-fit accelerator, attach a production endpoint, and scale it. We route to the cheapest accelerator that meets your latency and throughput targets, and we benchmark continuously. If the market can't beat renting yourself, we backstop with our own inference demand, so you’re never on the hook. 

You never pick the chip, cloud, or region; we handle routing, isolation, failover, metering, and billing across 4+ clouds and every major chip family. We're language- and framework-agnostic (no SDK lock-in), and you can drive everything from our CLI and REST API.

It's two-sided: if you run GPUs or other accelerators (a neocloud, a data center, a pay for an underutilized reserved fleet), you can connect that capacity to our rails and we'll bring the inference demand, turning hardware into a revenue-generating endpoint without building a platform yourself.

Traction

We're live in production, generating 100 billion tokens a week across 22 physical locations across Europe, APAC, North America, and the Middle East running on 8 different accelerator SKUs.  

Asks

We'd love to hear from you!

  • Run GPUs, or know someone who does? This is who we most want to talk to: companies running their own nodes for inference, data centers, reserved fleets, even a crypto or mining cluster looking to move into AI. Plug your capacity into our rails and we'll bring the demand. Reach us at founders@openrelay.inc.
  • Shipping inference and tired of picking chips? Bring a workload and we'll get you on the best available hardware today.
  • Feedback, intros, or just want to chat? Book a call with us, or email founders@openrelay.inc for anything else.
OpenRelay
Founded:2026
Batch:Summer 2026
Team Size:2
Status:
Active
Location:Seattle, WA
Primary Partner:Brad Flora