Favicon of mergekit

mergekit

A local LLM merging toolkit that combines model weights without extra training. Runs on CPU or GPU and supports Llama, Mistral and PyTorch checkpoints.

mergekit combines existing language models into a single model on your own hardware, without additional training or access to the original training data. It's for developers and researchers who want to combine fine-tuned capabilities or adjust the balance between model behaviors. The Python toolkit is open source under LGPL-3.0.

Merges can run entirely on CPU, or use GPU acceleration with as little as 8 GB of VRAM. It loads model weights only as needed, reducing memory demands when the input models don't fit in memory at once. Supported model families include Llama, Mistral, GPT-NeoX and StableLM.

Its methods include weighted averaging, SLERP, TIES, DARE and Arcee Fusion, giving users different ways to combine checkpoints and handle conflicting changes. It can also assemble models from selected layers, build mixture-of-experts models from dense models, and chain merges so one result becomes the input to another.

Beyond merging, mergekit can extract PEFT-compatible LoRA adapters from fine-tuned models. Tokenizer controls align vocabularies across inputs, while tokenizer transplantation supports draft models for speculative decoding. For work outside Hugging Face Transformers, it applies merge algorithms to raw PyTorch .pt and .safetensors checkpoints.

The toolkit runs locally. FrankensteinAI is a separate hosted service powered by mergekit, with a browser interface and a community gallery for sharing and comparing merged models.

Similar to mergekit