
Twinny is an AI coding assistant for VS Code that lets developers choose where their models run: on their own computer, a private server or a hosted API. It's for individuals and teams who want code suggestions and repository chat with control over where their code goes. The extension and team gateway are open source under the MIT license.
It supports Ollama, llama.cpp, LM Studio and any OpenAI-compatible server. You can use a local model for completion and a hosted model for chat. With local backends, it can work offline in an air-gapped network without an account or telemetry. Hosted providers receive the requests you send to them.
Completion uses surrounding code as context. Chat can draw on files, symbols, diagnostics, Git changes and terminal output, plus keyword and vector search across the workspace. Inline edits appear as diffs you can accept or reject by hunk. Reviews cover uncommitted changes, branches and GitHub pull requests; terminal assistance proposes commands for review and offers fixes when they fail.
The self-hosted twinny-server gateway gives a team shared access to its model servers and tracks usage by developer. It can also pool models running on teammates' computers, route requests to the least busy machine and switch machines if one disconnects. Prompts pass through the computer serving the model without being stored there. Gateway content recording is off by default.
Claim this page with an email at twinny.dev. Twinny gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Twinny?Promote it
Something wrong or outdated on this page?
27.5KUpdated 35 minutes agoMIT
VS Code · JetBrains#Code execution#Git integration#MCP
Kilo Code is an AI coding agent for developers who want help writing and understanding code in their editor or terminal. It runs in VS Code, JetBrains IDEs and the CLI, with cloud agents available for remote work. The project is open source under the MIT license and can be self-hosted.
6.9KUpdated 21 hours agoApache-2.0
#Agent Client Protocol#Human approval#MCP
4.8KUpdated 1 day ago
Docker · Web#Git integration#Human approval#LLM tracing
Agenta is an MIT-licensed, open source workspace for teams that want AI agents to handle tasks in chat and continue recurring work in the background. You can self-host it with Docker Compose or Helm to keep agents and workspace data on your infrastructure, or use Agenta Cloud as a hosted service.
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
1.4KUpdated 1 day ago
Docker · Web#Git integration#Multi-user access#OpenAI-compatible API
Kodus is an AI code review tool for engineering teams that want automated pull request feedback while choosing where the reviewer runs and which model it uses. Its open source core uses the AGPL license and can be self-hosted with Docker Compose or on Kubernetes. Kodus also offers a hosted cloud service that manages the infrastructure.
5.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend
codecompanion.nvim brings model chat, inline code edits, and coding agents into Neovim. It's for developers who want AI assistance inside their editor, with a choice between local LLMs through Ollama and cloud providers such as Anthropic, OpenAI, and Google Gemini. The plugin runs in Neovim; where model processing happens depends on the backend you connect.
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
llama-swap is a self-hosted proxy for people running several AI models on their own hardware. It starts the model server a request needs and swaps out another when necessary, so you don't have to keep every model loaded or manage separate API connections in your apps.