Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
Text embeddings let applications compare meaning for search, clustering and recommendations. Reranking scores documents against a query to reorder search results by relevance. Supported models include BAAI/bge-m3, jinaai/jina-embeddings-v3 and mixedbread-ai/mxbai-rerank-xsmall-v1, with models available from Hugging Face or local files.
Its media support covers CLIP and SigLIP for matching images with text, and LAION CLAP for matching audio with text. ColBERT and ColPali models support more detailed text and visual document retrieval. It can also run text classifiers for tasks such as sentiment analysis and emotion detection, and serve multiple models together.
The server exposes a REST API aligned with OpenAI's API specifications and an asynchronous Python interface. Docker deployment is supported. Inference can use CPUs, NVIDIA CUDA, AMD ROCm, Apple MPS or AWS Inferentia hardware, with PyTorch, ONNX/TensorRT and CTranslate2 backends. Dynamic batching groups requests to improve throughput under load.
Model inference runs on the machine hosting the server. Anonymous usage telemetry has an opt-out. Integrations include LangChain, Ragflow and KubeAI.
Claim this page and we'll verify you by hand. Infinity Embeddings gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Infinity Embeddings?Promote it
Something wrong or outdated on this page?
5.1KUpdated 1 week agoApache-2.0
macOS · Linux · Docker#Batch processing#Hugging Face integration#LLM tracing
Text Embeddings Inference is a self-hosted server for developers who need text embeddings for search and retrieval applications. It serves models through a REST API on your own hardware and can run offline once model weights are downloaded. The Rust project is open source under Apache 2.0.
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
8.4KUpdated 20 hours agoMIT
#Agent Skills#Batch processing#Guardrails
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.