
Text Embeddings Inference is a self-hosted server for developers who need text embeddings for search and retrieval applications. It serves models through a REST API on your own hardware and can run offline once model weights are downloaded. The Rust project is open source under Apache 2.0.
Model support includes Nomic, E5, Alibaba GTE and Qwen3 embeddings, alongside models such as sentence-transformers/all-mpnet-base-v2 and google/embeddinggemma-300m. It also serves rerankers that score how closely a passage matches a query, including BAAI/bge-reranker-base and Alibaba-NLP/gte-multilingual-reranker-base. Sequence classification covers tasks such as emotion detection, while SPLADE pooling produces sparse embeddings for BERT and DistilBERT models.
You can run it locally on a CPU or deploy it with Docker. GPU support covers NVIDIA hardware and AMD Instinct MI200 and MI300 series through ROCm. ARM64 CPU containers work on Apple Silicon, Ampere and Graviton hardware; Apple Silicon inference through Docker uses CPU cores rather than Metal or MPS. Access to private or gated Hugging Face models uses a token.
For services handling many requests, token-based dynamic batching groups work to use compute resources efficiently. TEI uses Flash Attention, Candle and cuBLASLt for inference, loads Safetensors weights, and doesn't require model graph compilation. Small Docker images and short startup times suit serverless deployments. OpenTelemetry tracing and Prometheus metrics provide visibility into the running service.
Claim this page and we'll verify you by hand. Text Embeddings Inference gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Text Embeddings Inference?Promote it
Something wrong or outdated on this page?
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
2.9KUpdated 6 months agoMIT
macOS · Docker#Batch processing#Hugging Face integration#Multimodal input
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
8.4KUpdated 20 hours agoMIT
#Agent Skills#Batch processing#Guardrails
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.