Favicon of Crawl4AI

Crawl4AI

Open-source web crawler that runs in Python or Docker, converts pages to Markdown and JSON, and offers a hosted API with MCP access.

Screenshot of Crawl4AI website

Crawl4AI is a self-hosted web crawler and scraper for developers building AI agents, retrieval-augmented generation (RAG) systems and data pipelines. It turns web pages into Markdown or structured JSON and runs as a Python library or a Docker server on your own hardware. The open-source code uses the Apache 2.0 license.

Its Markdown output preserves headings, tables and code blocks while filtering out menus, footers and other page clutter. It can filter content against a query and turn links into numbered references, so extracted text retains its sources.

Data extraction doesn't require an LLM. CSS, XPath and regex extraction work without one; model-based extraction supports open-source or hosted LLM providers and produces records that follow a defined JSON schema. It also accepts raw HTML and local files.

For sites that need a browser, Crawl4AI supports Chromium, Firefox and WebKit. It can process JavaScript pages, scroll through dynamically loaded content and reuse saved login sessions. Crawls can follow links across a site, resume after a crash and cache pages to avoid fetching them again. Screenshots, PDFs and media metadata are also available.

The cloud service runs the browsers and handles proxies, queues and retries on Crawl4AI's infrastructure. Self-hosting keeps crawler execution on your hardware; hosted requests go to the service. Cloud also adds web search and MCP access for clients such as Claude Code, Codex and Cursor, using an API key.

Similar to Crawl4AI