Favicon of Model2Vec

Model2Vec

An open-source Python package for local text embeddings on CPU, with MIT licensing and integrations for Sentence Transformers and LangChain.

Screenshot of Model2Vec website

Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.

The main tradeoff is speed and size versus accuracy. Model2Vec reports models up to 50 times smaller and CPU inference up to 500 times faster than the original model, with a small performance drop. Its base package has NumPy as its only major dependency, which keeps the software footprint small alongside the models.

You can use pretrained models or distill your own sentence transformer. Distillation needs a vocabulary and the source model, but no training dataset, and runs on CPU. Model2Vec also supports subword embeddings and fine-tuning classification models on top of its embeddings.

The pretrained range includes English general-purpose and retrieval models based on bge-base-en-v1.5, a multilingual model based on bge-m3, and a code model based on CodeRankEmbed. These give developers task-specific starting points without having to distill a model themselves.

Model2Vec integrates with Sentence Transformers and LangChain. Hugging Face Hub handles loading and sharing models, while the embedding models themselves run on your hardware.

Similar to Model2Vec