186.5KUpdated 1 day agoAGPL-3.0
Docker#Batch processing#MCP#Structured output
Firecrawl helps developers give AI agents and applications access to current web content. It searches for pages, extracts their contents, and returns data in forms an application can use. The project is open source under AGPL-3.0 and can run on a self-hosted server. Firecrawl also offers a hosted service that requires an account and API key and includes additional features. Reaching live websites requires an internet connection.
90.7KUpdated 1 week ago
Windows#Git integration#Knowledge graphs#MCP
MCP Reference Servers is a collection of locally run examples for developers building connections between AI applications and external tools or data. Maintained by the MCP steering group, the servers demonstrate the protocol and its SDKs. They're educational implementations, so developers should assess security requirements before using them in production.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
31.4KUpdated 5 days agoMIT
#MCP#Ollama integration
ScrapeGraphAI is an AI web scraping tool for developers who want to describe the data they need in plain language. Its open-source Python library runs on your own infrastructure under the MIT license. A separate managed API runs in ScrapeGraphAI's cloud.
84.5KUpdated 6 days agoApache-2.0
Docker#Structured output
Crawl4AI is a self-hosted web crawler and scraper for developers building AI agents, retrieval-augmented generation (RAG) systems and data pipelines. It turns web pages into Markdown or structured JSON and runs as a Python library or a Docker server on your own hardware. The open-source code uses the Apache 2.0 license.
25.5KUpdated 2 days agoMIT
#LLM tracing#Structured output
Stagehand is an open-source browser automation SDK for developers building AI agents that interact with websites and extract structured data. It can run with Chrome on your own machine or use Browserbase's cloud browsers. Local runs require Chrome. The project uses the MIT license and supports TypeScript, Python, and Go.
12.1KUpdated 4 months agoApache-2.0
Docker#Multimodal input#Structured output
Jina Reader turns web pages and documents into text that LLMs can use, with Markdown or JSON output. It's for developers building AI agents, search tools and systems that answer questions using retrieved documents. You can self-host the Apache 2.0 service code in Docker or use Jina's hosted API.
84.5KUpdated 1 day agoBSD-3-Clause
Docker#Guardrails#MCP#Tool calling
Scrapling is a Python web scraping framework for developers collecting website data or giving AI agents access to web pages. Its adaptive parser can find previously selected elements after a site's layout changes, reducing the need to repair extraction rules. It's open source under the BSD-3-Clause license and runs on your own machine or in Docker.
7.7KUpdated 2 days agoApache-2.0
macOS · Docker · Web
Steel is a browser API for developers building AI agents that interact with websites. It runs Chrome sessions locally or on a self-hosted server through Docker, so you can keep browser infrastructure on hardware you control. The code uses the Apache 2.0 license. Steel also offers a hosted service where browser sessions run in its cloud.