Scrapling is a Python web scraping framework for developers collecting website data or giving AI agents access to web pages. Its adaptive parser can find previously selected elements after a site's layout changes, reducing the need to repair extraction rules. It's open source under the BSD-3-Clause license and runs on your own machine or in Docker.
It handles ordinary HTTP requests and JavaScript-heavy pages through Playwright's Chromium or Google Chrome. Browser sessions retain cookies and state, and stealth fetching supports Cloudflare Turnstile and interstitial challenges. You can use local browsers or connect to browsers running on another host or with a managed provider. Fetching live websites needs network access; cached responses let you develop parsing logic without repeatedly contacting the site.
For larger collections, its crawler combines concurrent requests with multiple session types, proxy rotation and pause/resume. It adjusts request delays to each site's response speed and backs off when requests hit blocking or rate limits. Results can stream into a pipeline as they arrive, or export as JSON, JSONL, CSV or XML. Templates cover sitemap crawls, feeds and Shopify product collection.
The MCP server lets Claude, Cursor and other AI agents fetch pages, use browser sessions and capture screenshots. It can limit page content with CSS selectors and remove prompt-injection content before passing it to an agent. Scrapling also converts individual pages or whole sites into sanitized Markdown for RAG datasets without using an LLM. Existing Scrapy projects can use its parser on responses they already fetch.
Claim this page and we'll verify you by hand. Scrapling gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Scrapling?Promote it
Something wrong or outdated on this page?
186.5KUpdated 1 day agoAGPL-3.0
Docker#Batch processing#MCP#Structured output
Firecrawl helps developers give AI agents and applications access to current web content. It searches for pages, extracts their contents, and returns data in forms an application can use. The project is open source under AGPL-3.0 and can run on a self-hosted server. Firecrawl also offers a hosted service that requires an account and API key and includes additional features. Reaching live websites requires an internet connection.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
90.7KUpdated 1 week ago
Windows#Git integration#Knowledge graphs#MCP
MCP Reference Servers is a collection of locally run examples for developers building connections between AI applications and external tools or data. Maintained by the MCP steering group, the servers demonstrate the protocol and its SDKs. They're educational implementations, so developers should assess security requirements before using them in production.
31.4KUpdated 5 days agoMIT
#MCP#Ollama integration
ScrapeGraphAI is an AI web scraping tool for developers who want to describe the data they need in plain language. Its open-source Python library runs on your own infrastructure under the MIT license. A separate managed API runs in ScrapeGraphAI's cloud.
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.