Favicon of Llama Coder

Llama Coder

A local AI coding assistant for VS Code that uses Ollama on your computer or a server you control. Open source under MIT, with no telemetry.

Llama Coder is an open source VS Code extension for developers who want a self-hosted alternative to GitHub Copilot's code completion. It uses Ollama to run models on your own hardware, either on the computer you're coding on or on a separate machine. The extension has no telemetry or tracking.

Its focus is autocomplete within the editor. It supports programming languages and natural language text, along with Jupyter notebooks and remote files. You can pause suggestions and adjust the delay before they appear, so completion doesn't have to stay active throughout a coding session.

Model choices include Code Llama, Stable Code and DeepSeek, and the extension also lets you select a custom model. Stable Code is the default. Code Llama comes in several sizes and quantized variants, giving you choices that depend on the memory available on your machine.

You can keep inference on your workstation or send requests to an Ollama endpoint on a server you control. Remote inference supports bearer token authentication. This lets a dedicated GPU machine handle model execution while VS Code runs on another computer; the coding context used for those requests goes to that endpoint.

The stated minimum is 16 GB of RAM. Apple Silicon Macs and an NVIDIA RTX 4090 are recommended for performance, and Windows laptops with a suitable GPU are also supported. Larger models need more RAM or VRAM. The project uses the MIT license.

Similar to Llama Coder