{"id":111487,"title":"OneTriangle - The fastest, cheapest inference, powered by KV cache transfer","tagline":"Lightweight novel inference ","body":"**TL;DR:** OneTriangle is building the cheapest, fastest inference. We cut inference prefill costs by 20% and time-to-first-token by 40%, letting you get large-model quality at small-model cost.\n\nVideo: \u003chttps://youtu.be/HKe2EZUdNHo\u003e\n\n**Team**\n\nWe’re Hannah (CEO) and Medha (CTO), MIT CSAIL grads building with a team of 4 MIT engineers spanning ex-DeepMind and Jane Street, Physics and Astronomy Olympiad medalists, and NeurIPS/ICML authors. \n\n**The Problem**\n\nInference is expensive. \n\n• You’re forced into a bad tradeoff. You need the decoding power of large models, but you have the budget for small ones.\n\n• Prefill is the hidden tax. For long-context and agentic workloads, prefill dominates cost and latency, and everyone just eats it.\n\n**Our Solution**\n\nOneTriangle transfers the KV cache between models, something that has never been done in production. We prefill on a small model, strip the positional encoding, map the cache into the large model’s space, and restore relative importances. The result:\n\n• 20% lower prefill costs\n\n• 40% faster TTFT\n\n• Large-model quality at small-model cost\n\nWe beat out NVIDIA’s KV cache transfer paper by getting a lower KL divergence score in 2 weeks.\n\n![uploaded image](/media/?type=post\u0026id=111487\u0026key=user_uploads/2573695/e64727d6-a7b1-404a-bfbb-22ad2b941142)\n\n\\\nNo API changes, no quality cliff, and we’re upstreaming it into vLLM so the fastest path to cheap inference is the infrastructure you already run.\n\n**Why Now**\n\nInference has surpassed training as AI’s largest compute cost, and it’s accelerating with demand for agents that generate 10x the tokens a chat session does. Inference already eats up to 90% of total compute cost, over $600 billion is going into AI data centers this year alone. Every team pushing the frontier is bottlenecked by the same thing: serving costs.\n\n\\\n**The Future**\n\nOur bet is that model-mixing becomes the default serving pattern: small models handle prefill and routine tokens, large models handle the hard ones, and the cache moves freely between them. Whoever owns the transfer layer owns the economics of inference. We are the future of lightweight, fast inference.","slug":"T0B-onetriangle-the-fastest-cheapest-inference-powered-by-kv-cache-transfer","created_at":"2026-08-21T02:30:02.532Z","updated_at":"2026-09-19T08:02:00.424Z","total_vote_count":10,"url":"https://www.ycombinator.com/launches/T0B-onetriangle-the-fastest-cheapest-inference-powered-by-kv-cache-transfer","share_image_url":"//bookface-static.ycombinator.com/assets/ycdc/yc-og-image-c440a0ad1dacfb86eeeb343717479cc54d256614449b4ef719977a0a451f8bc8.png","company":{"id":31074,"name":"OneTriangle","slug":"onetriangle","url":"https://onetriangle.ai/","logo":"https://bookface-images.s3.amazonaws.com/small_logos/75197a50573558e4b8c16932b7db7c7c11f235a4.png","batch":"Summer 2026","industry":"B2B","tags":["Artificial Intelligence","Open Source","Infrastructure"],"search_path":"https://bookface.ycombinator.com/company/31074"}}