HomeCompaniesQuotient Labs
Quotient Labs

Use Claude Code at 47% less cost in one line of installation.

You should only be paying frontier model prices for top-shelf reasoning. But 88% of your token bill is just tool I/O. Quotient Labs runs on top of Claude Code to reduce token spend, compress context, and optimize tool calls while preserving output quality. Same models, same workflow, no changes to how you work. Benchmarked 47% less cost on real, multi-step developer workflows, with no quality loss.
Active Founders
Saahil Sundaresan
Saahil Sundaresan
Founder
Making coding agents cheaper @ Quotient Labs | Stanford CS & Linguistics | formerly at Apple, Amazon Feel free to reach out for anything! saahil@quotientlabs.com
Andrew Kuik
Andrew Kuik
Founder
Making coding agents cheaper @ Quotient Labs. Stanford CS; previously at AWS, Accenture Feel free to reach out for anything!
Company Launches
Quotient Labs: Stop overpaying Claude for tokens you don't need
See original launch post

TLDR: Tokenmaxxing is out. Tokenminning is in. Quotient Labs sits on top of Claude Code and reduces the amount of context you pay for—without changing your models, prompts, projects, or workflow. Across real, multi-step developer workflows, we’ve benchmarked 47% lower token cost with no measured quality loss. Get started today at https://www.quotientlabs.com/.

The Problem

Coding agents accumulate a lot of context.

Every file read, search result, test run, tool output, generated plan, and conversation turn gets added to the session. As the session grows, this history is repeatedly re-read (and re-billed) long after it has stopped being useful in its original form. The result? You’re overpaying Claude for tokens your agent never needed.

Fermat’s Last Token (our main product) sits on top of Claude Code and optimizes what enters and remains in your model’s context. We do this in four ways:

  1. Faster, more cost-efficient agent tools that batch operations and return compact, information-dense results.
  2. Bash output compression that preserves the useless parts of long command and test outputs while removing the bloat.
  3. Live text compression that removes the information-sparse and irrelevant tokens in conversation turns before they ever compound in cost. 
  4. Idle context optimization that compresses older conversation history as the prompt cache expires (this is more often than you think; only 5 minutes TTL) while preserving the freshest context verbatim.

The important part: you keep using Claude Code exactly as you do today.

Observability: See real-time estimated cost and savings down to the session and request level, as well as manage your organizations and users, in the user dashboard: https://console.quotientlabs.com/

Ready to go in 30 seconds. Install Fermat, sign in, and continue using Claude Code normally:

curl -fsSL https://downloads.quotientlabs.com/fermat/install.sh | bash && fermat login

Fermat is compatible with both CLI and the VS Code extension. Your existing projects, prompts, models, integrations, and interface are untouched.

You can bypass Fermat (if you really want) on a particular chat any time with claude --vanilla. You can also fully uninstall in one line with fermat uninstall (though we don’t recommend either :)).

Our ask

We’re looking for devs and engineering teams who use Claude Code heavily, especially on long-running, tool-intensive engineering tasks. 

We’d love to get feedback on where Fermat helps the most, where compression could be better, and any features or changes you’d like to see.

Get started: https://quotientlabs.com/

Or message us directly: founders@quotientlabs.com

Founders

Before Quotient Labs, we’ve been at Stanford (BS+MS CS/AI), industry (Apple, AWS, Accenture) and in academia (with multiple publications).

uploaded image

Quotient Labs
Founded:2025
Batch:Winter 2026
Team Size:2
Status:
Active
Location:San Francisco
Primary Partner:Ankit Gupta