Favicon of Infinity (InfiniFlow)

Infinity (InfiniFlow)

A self-hosted search database for LLM apps that combines vector and full-text retrieval. Runs in Docker or embedded in Python under Apache 2.0.

Screenshot of Infinity (InfiniFlow) website

Infinity is a self-hosted database for developers building search and retrieval-augmented generation (RAG) into LLM applications. It combines embedding search with full-text search and structured filters, so an application can retrieve relevant records through both meaning and exact terms.

Its search capabilities cover dense vectors, sparse vectors and tensors containing multiple vectors. Results can be reranked with reciprocal rank fusion (RRF), weighted sums or ColBERT. That gives developers ways to combine retrieval methods and adjust how the database ranks matches. Alongside embeddings, Infinity stores strings, numbers and other structured data that applications can use for filtering.

Infinity can run as a separate server in Docker or as a module embedded in Python. Its single-binary architecture has no dependencies, and developers can access it through Python and HTTP APIs. These deployment choices suit teams that want a shared search service as well as developers who want retrieval inside a Python application. Beyond question answering and conversational AI, its stated uses include recommendation systems, copilots and content generation.

The project is open source under Apache 2.0 and is written primarily in C++. The listed CPU requirement is x86_64 with AVX2 support. Supported environments include Linux and macOS, plus Windows through WSL or WSL2.

Similar to Infinity (InfiniFlow)