{"id":82563,"title":"Firecrawl 🔥: Open source crawling and scraping for AI-ready web data","tagline":"The easiest way to connect web data to your AI apps","body":"Hey everyone! We're Caleb, Nick, and Eric, the founders behind [Firecrawl](https://www.firecrawl.dev/) - an all-in-one developer platform for crawling \u0026 scraping web data for AI applications.\n\n**TLDR:** Firecrawl is an open source API that transforms any web data into a clean, LLM-ready format for RAG, agentic tasks, or training. Since launching in April we gained 8000 stars on GitHub ⭐️\n\n![uploaded image](/media/?type=post\u0026id=82563\u0026key=user_uploads/1045172/59b29329-b168-42b9-9642-0e11f99699de)\n\n\\\n**The Problem:** Our story began while building [Mendable.ai](http://Mendable.ai), one of the first managed RAG platforms used by companies like Coinbase, Snap, and MongoDB. We quickly discovered that web data was not only a popular source for AI applications but that its quality was crucial for successful deployments.\n\nBuilding a reliable stack that worked for almost any URL presented numerous challenges, and as we expanded, we encountered countless edge cases. While some great tools existed, none handled the entire process reliably. We envisioned an API that could take a URL, crawl its pages, and provide up-to-date, easy-to-use markdown.\n\nConversations with industry peers revealed that they were rebuilding similar infrastructure. This inspired us to create Firecrawl— a developer-friendly solution we wish we'd had from the start.\n\nWe launched a cloud offering over a weekend in April, and in just three months, we've garnered over 8,000 GitHub stars and empowered thousands of developers to transform web content into AI-ready data.\n\n**Our Solution:**\n\n![uploaded image](/media/?type=post\u0026id=82563\u0026key=user_uploads/1045172/14099b5d-73e8-4661-8ce4-65bb4623852f)\n\nOur open-source, developer-focused platform simplifies scraping \u0026 crawling for AI apps by handling:\n\n* Bypassing JavaScript rendering\n* Enriching metadata\n* Crawling without consistent sitemaps\n* Parallel scraping jobs\n* Hosting headless browsers and managing proxies\n* Bot blocking\n* Formatting LLM-friendly markdown\n\nWith Firecrawl, developers at companies like [Gamma](https://gamma.app/), [StackAI](https://www.stack-ai.com/), and [Zapier](https://zapier.com/) are delegating scraping to us so they can focus on their core tasks - be it RAG, agents, or data processing.\\\n\\\nNow scraping a whole website and retrieving the markdown is as simple as this:\n\n![uploaded image](/media/?type=post\u0026id=82563\u0026key=user_uploads/1045172/0c55112e-cd7a-4420-bd79-cbb07e4acaf3)\n\n**Our Asks:**\n\n* Give [Firecrawl](https://www.firecrawl.dev/) a try 🧑‍💻\n* Explore our [GitHub repository](https://github.com/mendableai/firecrawl) and consider leaving a star ⭐️\n* Drop me a line at [eric@firecrawl.com](mailto:eric@firecrawl.com) if you want to chat (we are super excited about collaborations) 🔥","slug":"LTf-firecrawl-open-source-crawling-and-scraping-for-ai-ready-web-data","created_at":"2024-07-30T17:47:16.926Z","updated_at":"2026-07-22T12:12:58.305Z","total_vote_count":59,"url":"https://www.ycombinator.com/launches/LTf-firecrawl-open-source-crawling-and-scraping-for-ai-ready-web-data","share_image_url":"https://www.ycombinator.com/media/?type=post\u0026id=82563\u0026key=user_uploads/1045172/14099b5d-73e8-4661-8ce4-65bb4623852f","company":{"id":27152,"name":"Firecrawl","slug":"firecrawl","url":"https://www.firecrawl.dev","logo":"https://bookface-images.s3.amazonaws.com/small_logos/fe35a47d1e4dd7a3c1a7dcbea86c08b5ef2d67e6.png","batch":"Summer 2022","industry":"B2B","tags":["Developer Tools","Open Source","AI"],"search_path":"https://bookface.ycombinator.com/company/27152"}}