OpenVINO Model Server and Continue: VS Code setup

Set up Qwen3-Coder in VS Code with Continue and OpenVINO Model Server using Windows binaries, an Intel GPU and an API on port 8000.

Player not loading? Watch on YouTube

Ezequiel Lanza demonstrates how to connect a local LLM served by OpenVINO Model Server to Continue in VS Code. The walkthrough uses Windows binaries; he also notes that Docker remains an option. The server exposes an OpenAI-compatible API that the editor extension can consume.

Setup starts with downloading and extracting the server package, then configuring environment variables. Lanza selects an optimized Qwen3-Coder model, which he says needs 12 GB of disk space. After creating a models folder, he uses one command to download the converted INT4 OpenVINO model from Hugging Face and start the server. The demonstrated configuration targets an Intel GPU, enables text generation and exposes the API on port 8000. It also assigns a model name for the client configuration.

In VS Code, he installs the open source Continue extension and edits its local assistant configuration with the model name and endpoint URL, using an example from the documentation. He describes support for chat, autocomplete and agentic coding tasks. The examples ask the model to explain OpenVINO Model Server and then inspect the models folder through filesystem tools in agentic mode.

The coding assistant runs on the presenter's own hardware. His positive assessment of speed is specific to this demonstration; the transcript provides no benchmark results or memory requirements.