Favicon of Infinity Embeddings

Infinity Embeddings

Self-hosted embedding and reranking API with MIT licensing, Hugging Face models, and CPU, NVIDIA, AMD and Apple MPS support.

Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.

Text embeddings let applications compare meaning for search, clustering and recommendations. Reranking scores documents against a query to reorder search results by relevance. Supported models include BAAI/bge-m3, jinaai/jina-embeddings-v3 and mixedbread-ai/mxbai-rerank-xsmall-v1, with models available from Hugging Face or local files.

Its media support covers CLIP and SigLIP for matching images with text, and LAION CLAP for matching audio with text. ColBERT and ColPali models support more detailed text and visual document retrieval. It can also run text classifiers for tasks such as sentiment analysis and emotion detection, and serve multiple models together.

The server exposes a REST API aligned with OpenAI's API specifications and an asynchronous Python interface. Docker deployment is supported. Inference can use CPUs, NVIDIA CUDA, AMD ROCm, Apple MPS or AWS Inferentia hardware, with PyTorch, ONNX/TensorRT and CTranslate2 backends. Dynamic batching groups requests to improve throughput under load.

Model inference runs on the machine hosting the server. Anonymous usage telemetry has an opt-out. Integrations include LangChain, Ragflow and KubeAI.

Similar to Infinity Embeddings