Favicon of Qwen2.5-Coder

Qwen2.5-Coder

Open-source coding LLMs for code generation, reasoning and fixes. Run the 32B instruction model locally with Transformers or self-host it with vLLM.

Screenshot of Qwen2.5-Coder website

Qwen2.5-Coder is a family of code-focused language models for developers who want to run a coding assistant on their own hardware or build a self-hosted code agent. It covers code generation, reasoning about code and fixing errors. The 32B instruction-tuned model is open source under Apache 2.0 and accepts requests in a chat format.

The family includes smaller models alongside the 32B model, giving developers a choice of model size for their deployment. Its training combines source code with text tied to code and synthetic data. It also retains mathematics and general language capabilities, so coding conversations can include explanations and related questions.

The 32B instruction model works with Hugging Face Transformers, and its weights come in Safetensors format. vLLM is a supported serving option for running it on your own server. Hugging Face also lists hosted inference through Nscale; that route uses a provider's infrastructure rather than local execution.

Long context is a notable feature for work involving substantial code inputs. The model supports up to 128K tokens with YaRN context extension, while its default configuration handles about 32K tokens. That larger window depends on support in the serving framework. Static YaRN scaling in vLLM can affect performance on shorter inputs.

Similar to Qwen2.5-Coder