Hi all 👋, We’re Michael and Charles, founders of Conifer.
TLDR: Conifer is one interface for both local and cloud models, available today as Juniper, our consumer-facing app. It routes each request to the most cost-effective option, beginning with your own hardware. Most requests are processed locally without API fees, reducing paid token volume by up to 80%.
📺 Launch Video: https://www.youtube.com/watch?v=QLVNISmtet0
❓ The problem: Model labs direct every request to the cloud, with no incentive to route your queries to a cheaper competitor or a free, on-device model. That may benefit them, but it’s unnecessarily costly for you. Whether it’s a simple typo fix or a complex architecture problem, the request is sent to a data center, exposing your data to third parties and increasing your monthly spend. This quickly adds up for teams running coding agents, customer support, and other high-volume workloads. When handling financials, patient records, or customer data, teams must also trust that a cloud provider’s security is as robust as promised.
🔨 What Conifer does: Conifer routes every request from the bottom up:
Tier 0: Your own hardware, with $0 in API fees
Tier 1: An efficient cloud model
Tier 2: A frontier model for the most demanding queries
The majority of requests are processed locally, ensuring that cloud pricing applies only when actually necessary, drastically reducing paid token volume. Additionally, Conifer consolidates disparate subscriptions, dashboards, and API keys within a single unified interface.
For sensitive workloads 🕵️, local-only mode disables cloud routing entirely, so all conversations and data remain completely on-device.
🏃♂️ Ensuring the local tier is fast enough: This routing only works if the local tier is reliable, so we built Conifer's inference engine from the ground up in Rust. On Apple Silicon, decode speeds reach up to 60% faster than llama.cpp.
On-device operation with Conifer is now the default starting point for inference, rather than a feature limited to hobbyist use.
Asks: