Favicon of Snowflake Arctic Embed

Snowflake Arctic Embed

Text embedding models for semantic search, licensed under Apache 2.0, with compact, large and long-context variants for document retrieval.

Screenshot of Snowflake Arctic Embed website

Snowflake Arctic Embed is a family of open-source text embedding models for developers building semantic search and document retrieval systems. It turns queries and documents into numerical representations that a search system can compare by meaning. The models use the Apache 2.0 license.

The family gives developers a choice between smaller models that prioritize latency and larger models that prioritize retrieval accuracy. snowflake-arctic-embed-xs builds on all-MiniLM-L6-v2, while the small, medium and large variants build on the intfloat/e5 family. The supplied MTEB/BEIR retrieval evaluations compare them with embedding models in similar size classes; these published comparisons describe the original releases. Later Arctic Embed 2.0 variants add multilingual retrieval.

For longer documents, snowflake-arctic-embed-m-long builds on nomic-ai/nomic-embed-text-v1-unsupervised. It accepts longer passages than the standard variants and can extend its context further with relative positional encoding. That makes it a relevant choice when a search workload needs to represent more text in each passage.

Training focuses on matching queries to relevant documents. The models learn from query-document pairs drawn from public datasets and proprietary web search data, then receive further training with relevant documents and difficult irrelevant examples. snowflake-arctic-embed-l is positioned as a replacement for closed embedding APIs, while the compact variants target workloads where inference latency and cost matter.

Similar to Snowflake Arctic Embed