
StarCoder2 is a family of code generation models for developers who want to run code completion on their own hardware or adapt a model to their code. It predicts code continuations rather than following conversational instructions, so it's suited to completion workflows rather than a chat-based coding assistant.
The family includes smaller and larger models, giving users a choice when balancing model size against available hardware. Training draws on programming languages in The Stack v2, alongside natural language material from Wikipedia, Arxiv, and GitHub issues. Its context window lets it use surrounding code when generating a continuation.
StarCoder2 works with Hugging Face Transformers and supports CPU, GPU, and multiple-GPU execution. It also works with Text Generation Inference for serving models. Quantization through bitsandbytes reduces the memory needed to load model weights, which matters if a full-precision model won't fit your hardware.
Developers can fine-tune the models on their own code or text datasets. The provided approach uses LoRA through PEFT with 4-bit quantization, reducing the resources needed for adaptation. BigCode-Evaluation-Harness supports evaluating StarCoder2 and its derivatives on code tasks. The project's Python repository uses the Apache 2.0 license.
Model weights are released under the BigCode OpenRAIL-M license, whose conditions are separate from the Apache-licensed code.
Claim this page with an email at bigcode-project.org. StarCoder2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find StarCoder2?Promote it
Something wrong or outdated on this page?
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
1.1KUpdated 2 years agoMIT
#Hugging Face integration#LoRA#Quantization
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
5.2KUpdated 4 days agoApache-2.0
Linux · Docker · Web#Hugging Face integration#LoRA#Quantization
960Updated 7 months agoApache-2.0
#Hugging Face integration#LoRA#Quantization
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
H2O LLM Studio is a self-hosted tool for teams that want to adapt language models to their own datasets without writing training code. Its browser interface brings training experiments, evaluation, and model testing into one place. The project is open source under Apache 2.0.
HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.