Openbenchmarks

We help humans & agents pick tools via Independent Benchmarks

Founding Engineer (AI + Backend)

$100K - $120K0.30% - 0.80%San Francisco, CA, US
Job type
Full-time
Role
Engineering, Backend
Experience
Any (new grads ok)
Visa
Will sponsor
Connect directly with founders of the best YC-funded startups.
Apply to role ›
Fenil Suchak
Fenil Suchak
CEO

About the role

Founding Engineer (AI + Backend)

About OpenBenchmarks

Agents are becoming first-class users and consumers of the internet. They research, evaluate, compare tools and increasingly make build-versus-buy decisions on behalf of people. Every company will need to get their products picked and used by agents.

Agents increasingly prefer open, independent and grounded benchmarks to make decisions.

Openbenchmarks is the evaluation infrastructure for agents - domain-specific, reproducible evaluations that help agents pick tools with confidence.

Our mission is to be the trusted evaluation layer for agents.

Founders previously led AI research and Infra teams at Oracle and Appfolio;

We started Openbenchmarks as an output of our research in the field of model behavior and how agents actually chose between different tools.

We're a team of researchers, engineers and work with the fastest growing AI first companies like Parallel, Firecrawl, Telnyx, TinyFish and more.

About the role

We're a team of researchers, engineers and theorists, in person in SF. You'll be building the benchmarks and the systems that run them.

Here are the broad themes that you’ll be working on

Benchmarks and evals design - designing domain-specific benchmarks from scratch: what to measure, how to ground it, and what makes a benchmark that an agent will continuously trust and pick from.

Continuously running large-scale autonomous eval systems - infrastructure that runs without a human in the loop: vendor APIs at scale, LLM-as-judge pipelines, scoring and metric computation, drift detection as models and products change underneath.

Research - understanding model behavior - why agents choose what they choose, and what survives their scrutiny. Adjacent to this you'll work on synthetic data generation, and problems like self-improvement code/software.

You'll build benchmarks & evals that the fastest-growing AI-first companies pay close attention to and produce continuous evals that their engineering teams care about.

If you want to do your life's work, reach out.

About the interview

  1. Intro Call
  2. Spend a day in office with the team in SF
  3. 2 day work trial
  4. Accepted

About Openbenchmarks

About OpenBenchmarks

Agents are becoming first-class users and consumers of the internet. They research, evaluate, compare tools and increasingly make build-versus-buy decisions on behalf of people. Every company will need to get their products picked and used by agents.

Agents increasingly prefer open, independent and grounded benchmarks to make decisions.

Openbenchmarks is the evaluation infrastructure for agents - domain-specific, reproducible evaluations that help agents pick tools with confidence.

Our mission is to be the trusted evaluation layer for agents.

Founders previously led AI research and Infra teams at Oracle and Appfolio;

We started Openbenchmarks as an output of our research in the field of model behavior and how agents actually chose between different tools.

We're a team of researchers, engineers and work with the fastest growing AI first companies like Parallel, Firecrawl, Telnyx, TinyFish and more.

Openbenchmarks
Founded:2024
Batch:F24
Team Size:4
Status:
Active
Location:San Francisco
Founders
Aditya Lahiri
Aditya Lahiri
CTO
Fenil Suchak
Fenil Suchak
CEO