Favicon of ScrapeGraphAI

ScrapeGraphAI

AI web scraping library and hosted API for structured data extraction. Run the MIT-licensed Python library yourself with Ollama or cloud LLM providers.

Screenshot of ScrapeGraphAI website

ScrapeGraphAI is an AI web scraping tool for developers who want to describe the data they need in plain language. Its open-source Python library runs on your own infrastructure under the MIT license. A separate managed API runs in ScrapeGraphAI's cloud.

The library extracts information from websites and local documents, including XML, HTML, JSON and Markdown. It uses LLMs and graph-based pipelines to turn a request into data extraction, with a standard pipeline for pulling information from a single page. This approach suits projects where writing page-specific selectors would add work.

You choose the model backend. The self-hosted library supports local LLMs through Ollama, alongside OpenAI, Groq, Gemini and Azure. With Ollama, model inference runs locally; choosing a cloud provider sends model requests to that service. The package collects anonymous usage metrics, which you can disable. Website scraping still involves accessing the target site.

The two offerings divide responsibility differently. With the library, you manage Playwright browser rendering, proxies, anti-bot handling and infrastructure. The hosted API manages those parts and the LLM, with Python and Node.js SDKs for calling the service.

Integrations include LangChain, LlamaIndex and Crew.ai for LLM applications, plus n8n, Dify and Zapier for workflow automation. ScrapeGraphAI also provides an MCP server.

Similar to ScrapeGraphAI