Favicon of FastEmbed

FastEmbed

A Python library for local embeddings and reranking, using ONNX Runtime with CPU or GPU support. Open source under Apache 2.0.

Screenshot of FastEmbed website

FastEmbed is a Python library that generates embeddings on your own hardware for semantic search and retrieval-augmented generation (RAG). It's for developers who need to turn text into searchable vectors without relying on a cloud embedding API. It can run on a CPU or use GPU acceleration, and its Apache 2.0 license makes it open source.

The library uses ONNX Runtime and quantized model weights, with few external dependencies and no requirement for PyTorch. That smaller dependency footprint also makes it suitable for serverless environments such as AWS Lambda, where the code runs in the cloud rather than on your machine. Data parallelism supports encoding large datasets.

FastEmbed covers several parts of a retrieval system:

  • Dense text embeddings, with Flag Embedding as the default and support for multilingual models.
  • Sparse text embeddings through SPLADE++.
  • Text cross encoders for reranking search results.

Developers can extend its dense embedding and cross encoder support with models outside the supplied selection. This gives projects room to choose a model suited to their text and retrieval task.

Qdrant maintains the library and integrates it into the Qdrant Python client. That integration connects document embedding with vector storage and text queries, including use with an in-process Qdrant database.

Similar to FastEmbed