
Snowflake Arctic Embed is a family of open-source text embedding models for developers building semantic search and document retrieval systems. It turns queries and documents into numerical representations that a search system can compare by meaning. The models use the Apache 2.0 license.
The family gives developers a choice between smaller models that prioritize latency and larger models that prioritize retrieval accuracy. snowflake-arctic-embed-xs builds on all-MiniLM-L6-v2, while the small, medium and large variants build on the intfloat/e5 family. The supplied MTEB/BEIR retrieval evaluations compare them with embedding models in similar size classes; these published comparisons describe the original releases. Later Arctic Embed 2.0 variants add multilingual retrieval.
For longer documents, snowflake-arctic-embed-m-long builds on nomic-ai/nomic-embed-text-v1-unsupervised. It accepts longer passages than the standard variants and can extend its context further with relative positional encoding. That makes it a relevant choice when a search workload needs to represent more text in each passage.
Training focuses on matching queries to relevant documents. The models learn from query-document pairs drawn from public datasets and proprietary web search data, then receive further training with relevant documents and difficult irrelevant examples. snowflake-arctic-embed-l is positioned as a replacement for closed embedding APIs, while the compact variants target workloads where inference latency and cost matter.
Claim this page with an email at snowflake.com. Snowflake Arctic Embed gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Snowflake Arctic Embed?Promote it
Something wrong or outdated on this page?
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
22.2KUpdated 1 week agoMIT
#Hugging Face integration#Multilingual
5.8KUpdated 2 days agoApache-2.0
#Hugging Face integration#Multilingual#Quantization
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
huggingface.coEmbedding and Reranker Models
#Batch processing#Hugging Face integration#Multilingual
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.
EmbeddingGemma is a text embedding model for developers building search and document features that run on phones, laptops or tablets. Based on Gemma 3, it converts text into numerical representations so applications can find related passages by meaning. Embeddings stay on your hardware, and the model works without an internet connection.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
GTE (General Text Embedding) is Alibaba’s family of downloadable models for representing text as vectors. Developers use these representations to compare queries with documents, cluster related text or supply retrieval components for larger applications.