Self-Hosted Embedding and Reranking Servers

Servers that expose embedding and reranker models over an API, like Text Embeddings Inference and Infinity Embeddings.

7 tools
Self-hosted AI application server with an OpenAI-compatible API, local Ollama and vLLM backends, document search and agent tool calling. MIT licensed.

8.4KUpdated 20 hours agoMIT

#Agent Skills#Batch processing#Guardrails

A self-hosted vector search engine for text and images, with built-in embedding generation, Docker deployment, and an Apache 2.0 license.

5KUpdated 6 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#RAG

A self-hosted LLM API server that runs ExLlamaV3 models on your hardware, with OpenAI-compatible endpoints and an AGPL-3.0 license.

1.4KUpdated 2 days agoAGPL-3.0

Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

A self-hosted text embedding server with a REST API, CPU and GPU support, and offline operation with downloaded model weights. Apache 2.0 licensed.

5.1KUpdated 1 week agoApache-2.0

macOS · Linux · Docker#Batch processing#Hugging Face integration#LLM tracing

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 1 day agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

Self-hosted embedding and reranking API with MIT licensing, Hugging Face models, and CPU, NVIDIA, AMD and Apple MPS support.

2.9KUpdated 6 months agoMIT

macOS · Docker#Batch processing#Hugging Face integration#Multimodal input

More in RAG and Vector Search