
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
Its main appeal is a common API across local and hosted backends. Ollama and vLLM handle local model inference; connections to OpenAI, Anthropic, Gemini, AWS Bedrock and Vertex AI send requests to those services. Where inference runs depends on the provider you choose. Applications can switch providers without changing their code.
OGX supports text and vision requests, streaming responses and embeddings. Its Responses API manages agent tool calls on the server, connects to MCP tools and searches uploaded files for retrieval-augmented generation (RAG). File and vector store APIs cover document storage and search, with backends including FAISS, SQLite-vec, Qdrant and PGVector.
Because applications communicate with OGX over HTTP, they're not tied to a Python framework. Existing OpenAI-compatible clients can use the server, and an Anthropic Messages API adapter supports the Anthropic SDK. The Responses API follows the Open Responses specification. OGX also manages prompt templates, processes documents for ingestion and supports batch requests.
Claim this page and we'll verify you by hand. OGX (formerly Llama Stack) gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OGX (formerly Llama Stack)?Promote it
Something wrong or outdated on this page?
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
10.5KUpdated 7 months agoApache-2.0
macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input
3.1KUpdated 2 days agoApache-2.0
macOS · Windows · Linux · Web#Code execution#MCP#Multi-agent workflows
20.1KUpdated 2 days agoMIT
macOS · Linux · Docker · Web#Code execution#llama.cpp backend#OpenAI-compatible API
157.6KUpdated 3 hours ago
Docker · Web#Code execution#MCP#OpenAI-compatible API
18.3KUpdated 24 hours agoMIT
macOS · Windows · Linux · Docker · Web#Human approval#Hybrid search#llama.cpp backend
aichat brings Ollama and cloud AI services into the same terminal interface for developers and people who work at the command line. It runs locally on macOS, Linux and Windows, with Android support through Termux. Model processing happens through the backend you choose: Ollama supports local models, while providers such as OpenAI, Claude and Gemini process requests in the cloud.
BotSharp is a self-hosted framework for .NET developers building AI agents into business applications. Written in C#, it runs on Windows, Linux and macOS and is open source software under Apache 2.0. Its plugin design lets teams choose their model provider, storage and interface while keeping agent coordination in the same framework.
DB-GPT is a self-hosted AI data assistant for teams analyzing business data and developers building data applications. It turns plain-language requests into SQL queries and Python analysis, then produces charts, dashboards, or HTML reports. You can run it on macOS or Linux, with Docker deployment also supported.
Dify is a source-available platform for teams building AI agents and apps on a visual canvas. Its Community Edition runs on your own server with Docker. Dify also offers a hosted cloud service, while Enterprise deployments can run in a VPC or on a self-hosted server. The Community Edition uses a custom Apache 2.0 derivative license.
DocsGPT is an MIT-licensed, open-source platform for teams that want AI search, assistants and agents over their own documents. It can run on your servers with local models, including fully air-gapped deployments where documents and questions stay inside your network. Answers include the source title and page number so readers can check the evidence.