Favicon of Text Embeddings Inference

Text Embeddings Inference

A self-hosted text embedding server with a REST API, CPU and GPU support, and offline operation with downloaded model weights. Apache 2.0 licensed.

Screenshot of Text Embeddings Inference website

Text Embeddings Inference is a self-hosted server for developers who need text embeddings for search and retrieval applications. It serves models through a REST API on your own hardware and can run offline once model weights are downloaded. The Rust project is open source under Apache 2.0.

Model support includes Nomic, E5, Alibaba GTE and Qwen3 embeddings, alongside models such as sentence-transformers/all-mpnet-base-v2 and google/embeddinggemma-300m. It also serves rerankers that score how closely a passage matches a query, including BAAI/bge-reranker-base and Alibaba-NLP/gte-multilingual-reranker-base. Sequence classification covers tasks such as emotion detection, while SPLADE pooling produces sparse embeddings for BERT and DistilBERT models.

You can run it locally on a CPU or deploy it with Docker. GPU support covers NVIDIA hardware and AMD Instinct MI200 and MI300 series through ROCm. ARM64 CPU containers work on Apple Silicon, Ampere and Graviton hardware; Apple Silicon inference through Docker uses CPU cores rather than Metal or MPS. Access to private or gated Hugging Face models uses a token.

For services handling many requests, token-based dynamic batching groups work to use compute resources efficiently. TEI uses Flash Attention, Candle and cuBLASLt for inference, loads Safetensors weights, and doesn't require model graph compilation. Small Docker images and short startup times suit serverless deployments. OpenTelemetry tracing and Prometheus metrics provide visibility into the running service.

Similar to Text Embeddings Inference