
Riften is an OpenAI- and Anthropic-compatible gateway that routes each LLM request to the lowest-cost model capable of completing it.
Connect your existing providers, change the base URL, and keep the product you already built. Riften works across customer-facing AI products and internal tools, including Claude Code and Codex.
Riften is already removing enterprise-scale model spend from production workloads.
Most AI products choose a model once and send nearly everything to it.
A documentation lookup, a support classification, and a codebase migration do not require the same intelligence. Yet they are routinely sent to the same frontier model at the same price.
For customer-facing products, that cost cuts directly into gross margin. For internal agents, it limits how broadly the company can deploy them.
Riften evaluates each request locally, applies the company’s cost, capability, latency, and reliability constraints, and selects among the models available to that organization.
Hard tasks can still reach frontier models. Routine work moves to lower-cost commercial or open-weight models. Teams can keep their provider accounts, pin a specific model when needed, and inspect the cost and routing decision behind every request.
The objective is not the cheapest token. It is the lowest expected cost of completing the work.
The router reveals which work is repeated. Outcomes reveal whether it was done correctly.
A passing test, an accepted patch, a resolved ticket, or a completed workflow can become part of a customer-specific evaluation. Riften uses those evaluations to improve routing, post-train open models around recurring work, and test those models against the APIs they are intended to replace.
Models earn production traffic through the same endpoint. Frontier APIs continue to handle novel work. Repeated workloads move onto models the company controls.
That creates a direct path from inference routing to evaluation, post-training, deployment, and the applications where the work happens.
The model labs are building intelligence every company can rent. Riften is building the path to intelligence every company can own.
Riften is founded by Carl Stauffer, a former Palantir Forward Deployed Engineer who shipped software into the purchasing and production systems of a nine-figure manufacturer. He was Palantir’s first Neurodivergent Fellow through Alex Karp’s program, conducted nuclear research at Jefferson Lab, studied computer science and applied mathematics at Johns Hopkins, and is an a16z Scout.
The rest of the team comes from Stanford Mathematics, Berkeley computer science, and Carnegie Mellon mathematics and computer science. Their work spans Oxford OATML, SPAR, CMU CyLab, Palo Alto Networks, Millennium, Amazon Search, NASA, UCSF Health, and industrial manufacturing systems.
They have built agent orchestration platforms, LLM evaluation and fine-tuning systems, cryptographic security software, production search infrastructure, CAD-matching systems, and autonomous QA for measuring agent performance. Their research includes an ICML acceptance and current work in review at NeurIPS.
The team also brings national-level recognition in mathematics and technical research spanning manufacturing, computational science, and new energy systems.
Riften combines enterprise deployment, frontier evaluation research, competitive mathematics, agent systems, security, and production ML in one team built for the full path from request to outcome to model.
We are looking for companies with meaningful LLM usage across customer-facing products or internal agents.
If most of your traffic still runs through one frontier model, send us:
We will show you what should remain on the frontier, what can move today, and whether Riften is worth testing.
Email us at info@riften.ai.