5.5KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API
Kokoro-FastAPI runs the Kokoro-82M speech model on your own machine or server and exposes an OpenAI-compatible speech API. It's for developers adding local text-to-speech to assistants, reading apps or audiobook workflows. Speech generation runs locally, and the API doesn't require an OpenAI account.
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
15.8KUpdated 11 hours agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
23.2KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
5.1KUpdated 1 week agoApache-2.0
macOS · Linux · Docker#Batch processing#Hugging Face integration#LLM tracing
Text Embeddings Inference is a self-hosted server for developers who need text embeddings for search and retrieval applications. It serves models through a REST API on your own hardware and can run offline once model weights are downloaded. The Rust project is open source under Apache 2.0.
2.5KUpdated 1 day ago
macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input
Roboflow Inference is a self-hosted computer vision server for teams building camera and image analysis systems. It runs on your own computer, server, or edge device and combines model predictions with workflows for tracking, counting, measuring, and responding to events. Roboflow also offers hosted servers and a Serverless Cloud API, where processing runs on its infrastructure.
5.8KUpdated 6 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
8.1KUpdated 3 days agoApache-2.0
#Batch processing#Distributed execution#Hugging Face integration
LMDeploy is an open-source toolkit for developers serving language and vision-language models on their own hardware. It combines model compression with inference and self-hosted APIs, so teams can use it for batch processing or as the model backend for an application. It uses the Apache 2.0 license.
14.8KUpdated 3 hours ago
Docker#Batch processing#Distributed execution#LoRA
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
3.7KUpdated 5 months agoMIT
Docker#OpenAI-compatible API#Streaming inference
Speaches is a self-hosted speech server for developers who want transcription, translation and speech generation on their own hardware. Its OpenAI-compatible API lets applications use local speech models through tools and SDKs built for OpenAI's API. The project is open source under the MIT license.
9.6KUpdated 7 hours agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
19.5KUpdated 11 hours agoApache-2.0
Docker#LoRA#Multimodal input#Prompt caching
KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
1.3KUpdated 2 days agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
3.3KUpdated 2 months agoMIT
Docker · Web#Multilingual#Speaker diarization#Voice activity detection
Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.
8.2KUpdated 1 hour ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
10.9KUpdated 1 day agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.
6.1KUpdated 6 hours agoApache-2.0
#Hugging Face integration#ONNX#OpenAI-compatible API
KServe is an open source platform for teams serving LLMs and predictive machine learning models on their own Kubernetes infrastructure. It puts both kinds of workloads under a common serving API, so teams can manage different model frameworks through the same platform. It uses the Apache 2.0 license.
2.9KUpdated 6 months agoMIT
macOS · Docker#Batch processing#Hugging Face integration#Multimodal input
Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
655Updated 3 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.