Favicon of OGX (formerly Llama Stack)

OGX (formerly Llama Stack)

Self-hosted AI application server with an OpenAI-compatible API, local Ollama and vLLM backends, document search and agent tool calling. MIT licensed.

Screenshot of OGX (formerly Llama Stack) website

OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.

Its main appeal is a common API across local and hosted backends. Ollama and vLLM handle local model inference; connections to OpenAI, Anthropic, Gemini, AWS Bedrock and Vertex AI send requests to those services. Where inference runs depends on the provider you choose. Applications can switch providers without changing their code.

OGX supports text and vision requests, streaming responses and embeddings. Its Responses API manages agent tool calls on the server, connects to MCP tools and searches uploaded files for retrieval-augmented generation (RAG). File and vector store APIs cover document storage and search, with backends including FAISS, SQLite-vec, Qdrant and PGVector.

Because applications communicate with OGX over HTTP, they're not tied to a Python framework. Existing OpenAI-compatible clients can use the server, and an Anthropic Messages API adapter supports the Anthropic SDK. The Responses API follows the Open Responses specification. OGX also manages prompt templates, processes documents for ingestion and supports batch requests.

Similar to OGX (formerly Llama Stack)