HomeCompaniesLitmus

Evals for humans

Litmus is building the most accurate framework for evaluating and benchmarking human capability, starting with software. AI will compound small differences in human capability into increasingly large differences in what people can accomplish, while making existing static benchmarks obsolete. Software is already there: AI can hill-climb any output-based evaluation, while the ability to direct it is becoming the defining advantage. Litmus applies the same approach we already use for model evals to humans – creating world-like environments, and inspecting trajectory instead of just output. Every knowledge industry will soon face the same problem. We build Litmus to tell you what humans are capable of. We're already helping build frontier technical teams at Mercor, Composio, Neo Scholars, and more.
Active Founders
Shaivi Rau
Shaivi Rau
Founder/CEO
CEO of Litmus
Elena Zhao
Elena Zhao
Founder/CTO
Evals for humans
Company Launches
Litmus - The industry's most accurate engineering interview
See original launch post

TL;DR: We evaluate candidates through async technical assessments built from your own context (repos, tickets, and job descriptions) and let them work in their own setup with all their tools while we track and evaluate their process and output. Engineering is changing; we help you hire for the role it’s becoming. Learn more at litmushiring.com.

https://www.youtube.com/watch?v=lusUktcmZGA

The problem
You're a month in with a new hire and they're just not getting it: struggling to ship independently, shipping AI-generated code they can't explain, clueless when production breaks. None of this showed up in your interview because your interview was never designed to catch it.

We all already agree LeetCode sucks. But work trials are time-draining, unscalable, and kind of awkward when you know within the first hour someone's not gonna make it. So most teams are stuck choosing between a broken screen and a process that doesn't scale.

What we built
Litmus fixes both. We turn your repos, tickets, and job descriptions into a full interview pipeline where every assessment is generated from scratch for your company. Not a generic template, your actual codebase, your real engineering problems, scoped as end-to-end features.

Candidates work in their own setup, with all their tools including AI because that's how engineering works now, and that's what you're hiring for. 

Every submission runs in a sandbox, so graders evaluate real code execution, not what looks good on paper, and grading is tuned to your company across tool fluency, process, and outcome. 

The result: a real read on how someone actually builds, not just how they interview.

  1. Generate Tailored Assessments

    uploaded image

  2. Candidates take Assessments

    uploaded image

  3. Evaluate at Scale

    uploaded image

  4. Deep Dive into Top Candidates

    uploaded image

Traction
3,000+ assessments evaluated. $60k ARR, growing fast. We've built the hiring system for Neo Scholars, their portfolio, and others. Let us bring our expertise to your team.

Who we are
We're Shaivi and Elena. As a SWE at Two Sigma and Meta, Elena watched people ace every interview and then fail to build anything real on the job. Shaivi worked at early-stage startups and in VC, sitting beside founders as they agonized over which engineers to trust with the product. Two sides of the same coin - the most important decision a company makes is its team, and current interviews still can't tell you who'll deliver. Litmus does.

Our ask
Scaling your eng team? We'll come to your office, set you up, and run Litmus on your next role. Book a demo at litmushiring.com. And if you know a team drowning in engineering interviews, send them our way.

YC Photos
Litmus
Founded:2025
Batch:Summer 2026
Team Size:3
Status:
Active
Location:New York City, NY
Primary Partner:Jared Friedman