XERJ

Local AI search engine that indexes code, documents and data. Runs offline on Linux, macOS and Windows, with an Elasticsearch-compatible API.

Screenshot of XERJ website

XERJ is a local search engine for AI agents that retrieves relevant code and documents instead of making an agent read whole files into its context. It's for developers building coding assistants, codebase Q&A or agents that need persistent memory. The open-source Rust engine runs on Linux, macOS and Windows, on a laptop or self-hosted server, under the Apache 2.0 license.

Code search is its main focus. It uses tree-sitter to index symbols and definitions, so an agent can retrieve a function, its signature and its file location. That supports reference coding: searching related open-source projects for an implementation before writing code. This approach is most useful for unfamiliar or private code; the reported benefit doesn't hold consistently for public libraries a model already knows.

Folder indexing also handles mixed documents and data, including PDF, DOCX, SQLite, CSV and logs. XERJ detects formats from file contents, infers field types and creates a searchable catalog. Indexing can resume after interruption.

Keyword and vector search share one engine, alongside agent memory and a knowledge graph with evidence attached to links. The built-in embedder works fully offline without an account or external embedding key. It uses lexical feature hashing, so retrieval depends on vocabulary overlap rather than neural understanding; a built-in neural encoder and an external embedding proxy are alternatives.

Its Elasticsearch-compatible HTTP API lets existing clients and Kibana connect directly.

The binary also provides an MCP stdio server for agent tool calls. It connects to an existing XERJ node; start that node separately before using MCP.

Similar to XERJ