{"id":114666,"title":"Real-SWE: A coding benchmark built from private company codebases","tagline":"Can coding agents can do real engineering work on codebases they've never seen?","body":"![uploaded image](/media/?type=post\u0026id=114666\u0026key=user_uploads/426892/94ad1344-e8cb-4757-8023-59f3ed3e6081)\n\n**TL;DR**\\\nSpecific Labs (YC F25) is launching Real-SWE, a benchmark that evaluates coding agents on real software engineering tasks that engineers performed, using private, out-of-distribution codebases. Public benchmarks are built on open source repos models have already seen. Real-SWE tests whether an agent can do the work inside a company it has never encountered. Leaderboard and task examples at \u003chttps://realswe.withspecific.com\u003e\\\n\\\n**Hi, we're Janak and Sid**\\\nOver the past year we've acquired and licensed operational data and codebases from real companies, and one question kept coming up with labs: how do coding agents actually perform on private code? Nobody had a clean way to measure it. So we built one.\n\n**The problem**\\\nEvery major coding benchmark is built on public repos. Models have trained on that code or on code that looks a lot like it. Scores keep climbing, but the number companies care about is different - can this agent land a fix in our billing system, our permissions layer, our customer data pipeline, without having seen any of it before?\n\n**What Real-SWE is**\\\nReal-SWE is a set of tasks pulled from real work engineers did on private company codebases. Examples include an app with 200K+ users, a fintech platform processing 100K+ bank statements, and enterprise sales tools.\n\nEach task gives the agent the codebase and the context an engineer would have had, then checks whether the fix actually works. One example: fix invoice billing so each business charges the right tax and exempt customers aren't taxed. The agent has to work out how the business handles tax, connect the tax provider, and keep invoices consistent.\\\n\\\n**What we found**\n\nAgents that look strong on public benchmarks struggle a lot more when the codebase is one they've never seen.\n\n![uploaded image](/media/?type=post\u0026id=114666\u0026key=user_uploads/426892/5c10421b-da9c-4447-bb76-5cf298a890c9)\n\n\\\n\\\n**The ask**\n\n* If you're at a lab and want your model evaluated, or want the full task set for training, email [janak@withspecific.com](mailto:janak@withspecific.com)\n* If you run a company and would be open to contributing a codebase (we handle anonymization and tasks stay private), reach out","slug":"TpS-real-swe-a-coding-benchmark-built-from-private-company-codebases","created_at":"2026-09-10T18:52:18.053Z","updated_at":"2026-09-19T09:00:27.990Z","total_vote_count":11,"url":"https://www.ycombinator.com/launches/TpS-real-swe-a-coding-benchmark-built-from-private-company-codebases","share_image_url":"https://www.ycombinator.com/media/?type=post\u0026id=114666\u0026key=user_uploads/426892/94ad1344-e8cb-4757-8023-59f3ed3e6081","company":{"id":31009,"name":"Specific Labs","slug":"specific-labs","url":"https://withspecific.com","logo":"https://bookface-images.s3.amazonaws.com/small_logos/f65a9129ea781a8d18fe1dc455b98be2430643a6.png","batch":"Fall 2025","industry":"B2B","tags":["Generative AI","SaaS","B2B"],"search_path":"https://bookface.ycombinator.com/company/31009"}}