Favicon of Scrapling

Scrapling

Python web scraping framework that runs locally or in Docker, with adaptive parsing, browser automation and an MCP server for AI agents. BSD-3-Clause licensed.

Scrapling is a Python web scraping framework for developers collecting website data or giving AI agents access to web pages. Its adaptive parser can find previously selected elements after a site's layout changes, reducing the need to repair extraction rules. It's open source under the BSD-3-Clause license and runs on your own machine or in Docker.

It handles ordinary HTTP requests and JavaScript-heavy pages through Playwright's Chromium or Google Chrome. Browser sessions retain cookies and state, and stealth fetching supports Cloudflare Turnstile and interstitial challenges. You can use local browsers or connect to browsers running on another host or with a managed provider. Fetching live websites needs network access; cached responses let you develop parsing logic without repeatedly contacting the site.

For larger collections, its crawler combines concurrent requests with multiple session types, proxy rotation and pause/resume. It adjusts request delays to each site's response speed and backs off when requests hit blocking or rate limits. Results can stream into a pipeline as they arrive, or export as JSON, JSONL, CSV or XML. Templates cover sitemap crawls, feeds and Shopify product collection.

The MCP server lets Claude, Cursor and other AI agents fetch pages, use browser sessions and capture screenshots. It can limit page content with CSS selectors and remove prompt-injection content before passing it to an agent. Scrapling also converts individual pages or whole sites into sanitized Markdown for RAG datasets without using an LLM. Existing Scrapy projects can use its parser on responses they already fetch.

Similar to Scrapling