
Crawl4AI is a self-hosted web crawler and scraper for developers building AI agents, retrieval-augmented generation (RAG) systems and data pipelines. It turns web pages into Markdown or structured JSON and runs as a Python library or a Docker server on your own hardware. The open-source code uses the Apache 2.0 license.
Its Markdown output preserves headings, tables and code blocks while filtering out menus, footers and other page clutter. It can filter content against a query and turn links into numbered references, so extracted text retains its sources.
Data extraction doesn't require an LLM. CSS, XPath and regex extraction work without one; model-based extraction supports open-source or hosted LLM providers and produces records that follow a defined JSON schema. It also accepts raw HTML and local files.
For sites that need a browser, Crawl4AI supports Chromium, Firefox and WebKit. It can process JavaScript pages, scroll through dynamically loaded content and reuse saved login sessions. Crawls can follow links across a site, resume after a crash and cache pages to avoid fetching them again. Screenshots, PDFs and media metadata are also available.
The cloud service runs the browsers and handles proxies, queues and retries on Crawl4AI's infrastructure. Self-hosting keeps crawler execution on your hardware; hosted requests go to the service. Cloud also adds web search and MCP access for clients such as Claude Code, Codex and Cursor, using an API key.
Claim this page with an email at crawl4ai.com. Crawl4AI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Crawl4AI?Promote it
Something wrong or outdated on this page?
12.1KUpdated 4 months agoApache-2.0
Docker#Multimodal input#Structured output
Jina Reader turns web pages and documents into text that LLMs can use, with Markdown or JSON output. It's for developers building AI agents, search tools and systems that answer questions using retrieved documents. You can self-host the Apache 2.0 service code in Docker or use Jina's hosted API.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
186.5KUpdated 1 day agoAGPL-3.0
Docker#Batch processing#MCP#Structured output
12KUpdated 1 week agoApache-2.0
Docker · Web#Batch processing#Human approval#Multi-agent workflows
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
Firecrawl helps developers give AI agents and applications access to current web content. It searches for pages, extracts their contents, and returns data in forms an application can use. The project is open source under AGPL-3.0 and can run on a self-hosted server. Firecrawl also offers a hosted service that requires an account and API key and includes additional features. Reaching live websites requires an internet connection.
Bisheng is an open source, self-hosted platform for teams building AI applications around business documents and processes. Its visual workflow editor combines automated tasks with human feedback, including intervention during multi-turn conversations. It's suited to document review, support ticket assistance and report generation that need more control than a single chatbot exchange.
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.