
Firecrawl helps developers give AI agents and applications access to current web content. It searches for pages, extracts their contents, and returns data in forms an application can use. The project is open source under AGPL-3.0 and can run on a self-hosted server. Firecrawl also offers a hosted service that requires an account and API key and includes additional features. Reaching live websites requires an internet connection.
Search can return the full content of results, so an agent can read a page without a separate scraping request. For a known URL, Firecrawl can produce clean Markdown, structured JSON, HTML, or a screenshot. It handles pages that rely on JavaScript and can extract content from web-hosted PDFs and DOCX files. Site-wide work is covered too: crawling collects pages, mapping finds URLs, and batch scraping processes multiple URLs.
The hosted Firecrawl service also offers Interact for clicking, scrolling and entering text on pages, plus an Agent that can gather information without starting URLs. Its separate MCP server can point at a self-hosted Firecrawl API, and the CLI supports a custom API URL. Python and Node.js SDKs are available. AI-backed features on a self-hosted installation need a configured model provider.
Claim this page with an email at firecrawl.dev. Firecrawl gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Firecrawl?Promote it
Something wrong or outdated on this page?
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
84.5KUpdated 1 day agoBSD-3-Clause
Docker#Guardrails#MCP#Tool calling
90.7KUpdated 1 week ago
Windows#Git integration#Knowledge graphs#MCP
MCP Reference Servers is a collection of locally run examples for developers building connections between AI applications and external tools or data. Maintained by the MCP steering group, the servers demonstrate the protocol and its SDKs. They're educational implementations, so developers should assess security requirements before using them in production.
31.4KUpdated 5 days agoMIT
#MCP#Ollama integration
ScrapeGraphAI is an AI web scraping tool for developers who want to describe the data they need in plain language. Its open-source Python library runs on your own infrastructure under the MIT license. A separate managed API runs in ScrapeGraphAI's cloud.
1.1KUpdated 7 hours agoMIT
macOS · Windows · Linux#Agent Skills#MCP#Persistent memory
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Scrapling is a Python web scraping framework for developers collecting website data or giving AI agents access to web pages. Its adaptive parser can find previously selected elements after a site's layout changes, reducing the need to repair extraction rules. It's open source under the BSD-3-Clause license and runs on your own machine or in Docker.
deja-vu gives coding assistants a shared memory of work already recorded on your machine. It searches sessions from before you installed it, so developers can recover an old fix or carry context between Claude Code, Codex CLI, Cursor and opencode without starting a separate collection of notes.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.