Favicon of Firecrawl

Firecrawl

An open source web scraping API you can self-host or use as a hosted service to turn websites into Markdown, JSON, HTML, or screenshots.

Screenshot of Firecrawl website

Firecrawl helps developers give AI agents and applications access to current web content. It searches for pages, extracts their contents, and returns data in forms an application can use. The project is open source under AGPL-3.0 and can run on a self-hosted server. Firecrawl also offers a hosted service that requires an account and API key and includes additional features. Reaching live websites requires an internet connection.

Search can return the full content of results, so an agent can read a page without a separate scraping request. For a known URL, Firecrawl can produce clean Markdown, structured JSON, HTML, or a screenshot. It handles pages that rely on JavaScript and can extract content from web-hosted PDFs and DOCX files. Site-wide work is covered too: crawling collects pages, mapping finds URLs, and batch scraping processes multiple URLs.

The hosted Firecrawl service also offers Interact for clicking, scrolling and entering text on pages, plus an Agent that can gather information without starting URLs. Its separate MCP server can point at a self-hosted Firecrawl API, and the CLI supports a custom API URL. Python and Node.js SDKs are available. AI-backed features on a self-hosted installation need a configured model provider.

Similar to Firecrawl