Favicon of Nomic Embed Text v1.5

Nomic Embed Text v1.5

A local text embedding model for semantic search and RAG, with adjustable vector sizes, Apache 2.0 licensing, and support for Sentence Transformers.

Screenshot of Nomic Embed Text v1.5 website

Nomic Embed Text v1.5 is an English text embedding model for developers building semantic search, document retrieval, and RAG applications on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its main distinction is adjustable embedding size: you can use smaller vectors when storage matters, with a tradeoff in retrieval quality.

That flexibility comes from Matryoshka Representation Learning, which trains the model to retain useful information in shortened embeddings. Long-context support also lets it process longer documents and passages, with sequences of up to 8,192 tokens.

The model distinguishes between documents and search queries, so retrieval applications can represent each according to its role. It also supports embeddings for grouping texts by topic, finding semantic duplicates, and supplying features to a classification model.

You can run the model locally through Sentence Transformers, Hugging Face Transformers, or Transformers.js, and serve it through Infinity in Docker. Local inference processes text on your own hardware; the separate Nomic Embedding API processes requests through a hosted service. The model uses the Apache 2.0 license, and Nomic releases its training data in full.

For applications that search across text and images, nomic-embed-vision-v1.5 shares the text model's embedding space, allowing text embeddings to work alongside image embeddings.

Similar to Nomic Embed Text v1.5