New API is a self-hosted AI gateway for developers and teams that want several model providers behind one service. It builds on One API and converts between OpenAI Chat Completions, Responses, Anthropic Messages and Gemini formats, so apps and agents can switch providers without changing each client's connection settings.
The gateway runs on your own machine or server through Docker, with Linux containers for AMD64 and ARM64. Model requests go to the connected upstream services. Cloud providers require internet access and authorized API keys or accounts; hosting the gateway yourself doesn't keep those requests on your hardware. Supported providers include OpenAI, Anthropic, Google Gemini, DeepSeek and Qwen, alongside Azure OpenAI, AWS Bedrock and Vertex AI.
Routing includes provider priorities, weighted distribution and retries. Teams can control access through user groups, permissions and API key restrictions, with OAuth/OIDC, passkeys and two-factor authentication for sign-in. The web console brings model management, usage and audit logs, cost accounting and a model playground into one interface.
Beyond text chat, it handles image and audio APIs, embeddings and reranking. JavaScript plugins extend image, video and other asynchronous tasks. Streaming, tool use, reasoning, multimodal inputs and WebSocket connections depend on the upstream model and conversion path; some protocol-specific fields don't map exactly. New API is open source under AGPL-3.0, with additional attribution terms for modified versions.
Claim this page and we'll verify you by hand. New API gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find New API?Promote it
Something wrong or outdated on this page?
37.5KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#LLM tracing#MCP#Multimodal input
Claude Code Router is an open-source local model gateway for developers who use coding agents and want to manage their model providers in one place. It runs on macOS, Windows and Linux, with Docker and a CLI with a browser interface also available. The project uses the MIT license.
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
37KUpdated 2 years agoMIT
Linux · Docker · Web#Multi-user access#Ollama integration#OpenAI-compatible API
5.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
10.5KUpdated 7 months agoApache-2.0
macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
One API is a self-hosted LLM API gateway for developers and teams that want to share model access across apps or users. It puts cloud providers and Ollama behind an OpenAI-compatible API, so clients can use one endpoint across different backends. It's open source under MIT and runs on your own server as a single executable or in Docker.
llama-swap is a self-hosted proxy for people running several AI models on their own hardware. It starts the model server a request needs and swaps out another when necessary, so you don't have to keep every model loaded or manage separate API connections in your apps.
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
aichat brings Ollama and cloud AI services into the same terminal interface for developers and people who work at the command line. It runs locally on macOS, Linux and Windows, with Android support through Termux. Model processing happens through the backend you choose: Ollama supports local models, while providers such as OpenAI, Claude and Gemini process requests in the cloud.