Favicon of StarCoder2

StarCoder2

Code completion models that run locally on CPU or GPUs, with Hugging Face Transformers support and LoRA fine-tuning.

Screenshot of StarCoder2 website

StarCoder2 is a family of code generation models for developers who want to run code completion on their own hardware or adapt a model to their code. It predicts code continuations rather than following conversational instructions, so it's suited to completion workflows rather than a chat-based coding assistant.

The family includes smaller and larger models, giving users a choice when balancing model size against available hardware. Training draws on programming languages in The Stack v2, alongside natural language material from Wikipedia, Arxiv, and GitHub issues. Its context window lets it use surrounding code when generating a continuation.

StarCoder2 works with Hugging Face Transformers and supports CPU, GPU, and multiple-GPU execution. It also works with Text Generation Inference for serving models. Quantization through bitsandbytes reduces the memory needed to load model weights, which matters if a full-precision model won't fit your hardware.

Developers can fine-tune the models on their own code or text datasets. The provided approach uses LoRA through PEFT with 4-bit quantization, reducing the resources needed for adaptation. BigCode-Evaluation-Harness supports evaluating StarCoder2 and its derivatives on code tasks. The project's Python repository uses the Apache 2.0 license.

Model weights are released under the BigCode OpenRAIL-M license, whose conditions are separate from the Apache-licensed code.

Similar to StarCoder2