HomeCompaniesSpecific Labs
Specific Labs

Real World Environments and Data for AI

Specific builds enterprise environments and datasets representative of work inside real-world businesses. The next frontier of AGI is bottlenecked by data that is hard to access, and we make that easier.
Active Founders
Janak Sunil
Janak Sunil
Founder/CEO
Co-Founder and CEO @ Specific Prev @ Coinbase
Siddhant Paliwal
Siddhant Paliwal
Founder/CTO
Co-Founder and CTO @ Specific Prev @ Third Chair (YC X25), Intel. Founded my first company at 15, invited to the UN.
Company Launches
Real-SWE: A coding benchmark built from private company codebases
See original launch post

uploaded image

TL;DR
Specific Labs (YC F25) is launching Real-SWE, a benchmark that evaluates coding agents on real software engineering tasks that engineers performed, using private, out-of-distribution codebases. Public benchmarks are built on open source repos models have already seen. Real-SWE tests whether an agent can do the work inside a company it has never encountered. Leaderboard and task examples at https://realswe.withspecific.com

Hi, we're Janak and Sid
Over the past year we've acquired and licensed operational data and codebases from real companies, and one question kept coming up with labs: how do coding agents actually perform on private code? Nobody had a clean way to measure it. So we built one.

The problem
Every major coding benchmark is built on public repos. Models have trained on that code or on code that looks a lot like it. Scores keep climbing, but the number companies care about is different - can this agent land a fix in our billing system, our permissions layer, our customer data pipeline, without having seen any of it before?

What Real-SWE is
Real-SWE is a set of tasks pulled from real work engineers did on private company codebases. Examples include an app with 200K+ users, a fintech platform processing 100K+ bank statements, and enterprise sales tools.

Each task gives the agent the codebase and the context an engineer would have had, then checks whether the fix actually works. One example: fix invoice billing so each business charges the right tax and exempt customers aren't taxed. The agent has to work out how the business handles tax, connect the tax provider, and keep invoices consistent.

What we found

Agents that look strong on public benchmarks struggle a lot more when the codebase is one they've never seen.

uploaded image



The ask

  • If you're at a lab and want your model evaluated, or want the full task set for training, email janak@withspecific.com
  • If you run a company and would be open to contributing a codebase (we handle anonymization and tasks stay private), reach out
Previous Launches
Track and influence how often you appear on ChatGPT, Claude Code, Perplexity, and more
Specific Labs
Founded:2025
Batch:Fall 2025
Team Size:6
Status:
Active
Location:San Francisco
Primary Partner:Pete Koomen