Favicon of Jina Reader

Jina Reader

Self-hosted web extraction API converts pages and PDFs to Markdown for LLMs. Apache 2.0 service code runs in Docker; a hosted API is also available.

Screenshot of Jina Reader website

Jina Reader turns web pages and documents into text that LLMs can use, with Markdown or JSON output. It's for developers building AI agents, search tools and systems that answer questions using retrieved documents. You can self-host the Apache 2.0 service code in Docker or use Jina's hosted API.

Reader removes page clutter such as menus and ads while keeping the main content. Its browser renderer handles pages that need JavaScript, and extraction controls let you focus on specific sections. It reads PDFs and Word, Excel and PowerPoint files; image captions give text-only models some context about visuals. ReaderLM-v2 can extract structured fields using a JSON schema or written instructions.

The hosted service also provides web search with result URLs and page content. An MCP server connects Jina's APIs to LLM tools. Hosted requests run on Jina's servers, and a privacy setting prevents request caching and logging.

The service code and reader models have different licenses. ReaderLM-v2 and jina-vlm use CC-BY-NC 4.0, which permits noncommercial use; commercial production use requires a separate license. On-premises model containers run offline under the Jina On-Prem commercial license and make no outbound connections. Fetching public web pages still needs network access. The self-hosted service can cache fetched pages in S3-compatible storage.

Similar to Jina Reader