
Jina Reader turns web pages and documents into text that LLMs can use, with Markdown or JSON output. It's for developers building AI agents, search tools and systems that answer questions using retrieved documents. You can self-host the Apache 2.0 service code in Docker or use Jina's hosted API.
Reader removes page clutter such as menus and ads while keeping the main content. Its browser renderer handles pages that need JavaScript, and extraction controls let you focus on specific sections. It reads PDFs and Word, Excel and PowerPoint files; image captions give text-only models some context about visuals. ReaderLM-v2 can extract structured fields using a JSON schema or written instructions.
The hosted service also provides web search with result URLs and page content. An MCP server connects Jina's APIs to LLM tools. Hosted requests run on Jina's servers, and a privacy setting prevents request caching and logging.
The service code and reader models have different licenses. ReaderLM-v2 and jina-vlm use CC-BY-NC 4.0, which permits noncommercial use; commercial production use requires a separate license. On-premises model containers run offline under the Jina On-Prem commercial license and make no outbound connections. Fetching public web pages still needs network access. The self-hosted service can cache fetched pages in S3-compatible storage.
Claim this page with an email at jina.ai. Jina Reader gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Jina Reader?Promote it
Something wrong or outdated on this page?
84.5KUpdated 6 days agoApache-2.0
Docker#Structured output
Crawl4AI is a self-hosted web crawler and scraper for developers building AI agents, retrieval-augmented generation (RAG) systems and data pipelines. It turns web pages into Markdown or structured JSON and runs as a Python library or a Docker server on your own hardware. The open-source code uses the Apache 2.0 license.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
186.5KUpdated 1 day agoAGPL-3.0
Docker#Batch processing#MCP#Structured output
39.9KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#Knowledge graphs#LLM tracing#Multimodal input
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Firecrawl helps developers give AI agents and applications access to current web content. It searches for pages, extracts their contents, and returns data in forms an application can use. The project is open source under AGPL-3.0 and can run on a self-hosted server. Firecrawl also offers a hosted service that requires an account and API key and includes additional features. Reaching live websites requires an internet connection.
LightRAG combines knowledge graphs with vector search to answer questions across a document collection. It's a self-hosted Python framework for developers building document assistants, particularly where answers depend on relationships between facts in different files, such as legal or financial material.