{"id":81814,"title":"🚀 Hamming - Make your RAG \u0026 AI agents reliable","tagline":"The only end-to-end AI development platform you need: prompt management, evals, observability","body":"👋 [@Sumanyu Sharma](https://www.linkedin.com/in/sumanyusharma) and [@Marius Buleandra](https://www.linkedin.com/in/mariusbuleandra) from [@Hamming AI](https://www.ycombinator.com/companies/hamming-ai)\n\n**TLDR:** Are you struggling to make your RAG \u0026 AI agents reliable? We're launching our [AI Optimization Platform](https://app.hamming.ai) to help eng teams speed up iteration velocity, root-cause bad AI outputs, and prevent regressions.\n\n🌟 [Click here to try our AI Optimization Platform](https://app.hamming.ai) 🌟\n\n## Our thesis: Experimentation drives reliability\n\nPreviously, [Marius](https://www.linkedin.com/in/mariusbuleandra/) and I ran growth-focused eng and data teams at Tesla, Citizen, Spell \u0026 Anduril. We learned that running experiments is the best way to move a metric. More experiments = more growth.\n\nLast year, we shipped 10+ RAG \u0026 AI agents to production. We found the same pattern holds when building reliable AI products. More experiments = more reliability = more retention for your AI products.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/96653f00-344f-46b7-92da-7530f0302ae0)\n\n## Problem: Making RAG and AI agents reliable feels like whack-a-mole\n\nHere's the workflow most teams follow:\n\n1. **Tweak** your RAG or AI agents by indexing new documents, adding new tools, changing the prompts, models or other business logic.\n2. **Eyeball** how well your changes improved a handful of examples you wanted to fix. Often ad-hoc and slow.\n3. **Ship** the changes if they worked.\n4. **Detect** regressions when users complain of things breaking in production.\n5. **Repeat** steps 1 to 4 until you get tired.\n\nSteps 2 and 4 are often the slowest \u0026 most painful parts of the feedback loop. This is what we tackle.\n\n## Our take: Use LLMs as judges to speed up iteration velocity\n\nWe use LLMs to score the outputs of other LLMs. This is the fastest way to speed up the feedback loop.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/1e6dff78-03d8-4add-9a74-7382a9733147)\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/46cb1626-03be-4b9e-9218-c86aa4794b1f)\n\n### Flag errors in production before customers notice\n\nWe go beyond passive LLM / trace-level monitoring. We actively **score your production outputs** in real time and **flag cases** the team needs to double-click on. This helps eng teams quickly prioritize cases they need to fix.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/69edba93-dfb5-448d-8215-c3f3cb81402b)\n\n### Test new changes quickly while developing and prevent regressions from reaching users\n\nWe make offline evaluations easy, so you can change your system and get feedback in minutes.\n\n**Eval-driven prompt iteration**\n\nRapidly iterate with new prompts and models in our prompt playground with first-class support for function-calling. Run evals so you know your changes are improving things.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/89972e98-9709-47b1-b887-7c02a1af51a1)\n\n**Deploy prompt changes without code change**\n\nWe keep track of all prompt versions and update your prompts on the fly without needing a code change.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/94c56689-bac3-4e05-bb13-22234edce861)\n\n**Easily create golden datasets**\n\nOffline evaluations are bottlenecked on a high-quality golden dataset of input/output pairs. We support converting production traces to dataset examples in one click.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/4d849d21-b8e6-4beb-8617-215296de422a)\n\n**Diagnose between retrieval, reasoning or function-calling errors quickly**\n\nDifferentiating between retrieval, reasoning, and function-calling errors is time-consuming. We score each retrieved context on metrics like hallucination, recall, and precision to help you prioritize your eng efforts where it matters.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/afb539a0-1a67-4050-8733-387b03a536e6)\n\n**Override AI scores**\n\nSometimes our AI scores disagree with your definition of \"good\". We make it easy to override our scores with your preferences. Our AI scorer learns from your feedback.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/8227f0e2-e20d-4d2b-8ab6-aa413ec7efe6)\n\n## Meet the team\n\n[Sumanyu](https://www.linkedin.com/in/sumanyusharma/) previously helped Citizen (safety app; backed by Founders Fund, Sequoia, 8VC) grow its users by 4X and grew an AI-powered sales program to $100s of millions in revenue/year at Tesla.\n\n[Marius](https://www.linkedin.com/in/mariusbuleandra/) previously ran data infrastructure @ Anduril, drove user growth at Citizen with Sumanyu and was a founding engineer @ Spell (MLOps startup acquired by Reddit).\n\n![Sumanyu \u0026 Marius](https://bookface.ycombinator.com/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/750619c5-790d-4ba7-a4a4-e60c8f520506)\n\n## Our offer\n\nWe previously launched [Prompt Optimizer](https://www.ycombinator.com/launches/L4V-hamming-let-ai-optimize-your-prompts-free-for-7-days) on launch YC, which saves 80% of manual prompt engineering effort. In this launch, we show how teams use Hamming to build reliable RAG and AI agents.\n\n![uploaded image](/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/234292dd-27c0-4073-8082-7fc0340f1221)\n\n**Free 1:1 Debug session.** Struggling with making your RAG/agents reliable? We're offering a complementary 1:1 RAG/agent debugging session. Book time with us [here.](https://calendly.com/sumanyusharma/1-1-debug-session)\n\nQuestions? Email us [here](mailto:sumanyu@hamming.ai) or chat with us [here](https://calendly.com/sumanyusharma/30min).","slug":"LHa-hamming-make-your-rag-ai-agents-reliable","created_at":"2024-06-28T01:28:56.750Z","updated_at":"2026-07-22T05:09:19.588Z","total_vote_count":118,"url":"https://www.ycombinator.com/launches/LHa-hamming-make-your-rag-ai-agents-reliable","share_image_url":"https://www.ycombinator.com/media/?type=post\u0026id=81814\u0026key=user_uploads/497858/234292dd-27c0-4073-8082-7fc0340f1221","company":{"id":29610,"name":"Hamming AI","slug":"hamming-ai","url":"https://hamming.ai/","logo":"https://bookface-images.s3.amazonaws.com/small_logos/63f48bd5017efb50e1fd3bcb1bd32b4b6a0147e3.png","batch":"Summer 2024","industry":"B2B","tags":[],"search_path":"https://bookface.ycombinator.com/company/29610"}}