
Nomic Embed Text v1.5 is an English text embedding model for developers building semantic search, document retrieval, and RAG applications on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its main distinction is adjustable embedding size: you can use smaller vectors when storage matters, with a tradeoff in retrieval quality.
That flexibility comes from Matryoshka Representation Learning, which trains the model to retain useful information in shortened embeddings. Long-context support also lets it process longer documents and passages, with sequences of up to 8,192 tokens.
The model distinguishes between documents and search queries, so retrieval applications can represent each according to its role. It also supports embeddings for grouping texts by topic, finding semantic duplicates, and supplying features to a classification model.
You can run the model locally through Sentence Transformers, Hugging Face Transformers, or Transformers.js, and serve it through Infinity in Docker. Local inference processes text on your own hardware; the separate Nomic Embedding API processes requests through a hosted service. The model uses the Apache 2.0 license, and Nomic releases its training data in full.
For applications that search across text and images, nomic-embed-vision-v1.5 shares the text model's embedding space, allowing text embeddings to work alongside image embeddings.
Claim this page and we'll verify you by hand. Nomic Embed Text v1.5 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Nomic Embed Text v1.5?Promote it
Something wrong or outdated on this page?
2KUpdated 1 year ago
Docker#Batch processing#Hugging Face integration#Multilingual
Qwen3-Embedding is a family of text embedding models for developers building search and document analysis on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its multilingual support covers languages including English, Chinese, Arabic and Ukrainian, as well as programming languages.
jina.aiEmbedding and Reranker Models
Docker#GGUF#LoRA#MLX
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
huggingface.coEmbedding and Reranker Models
#Batch processing#Hugging Face integration#Multilingual
34.4KUpdated 6 months agoMIT
#Batch processing#Multimodal input
CLIP is an open-source image and text model that lets developers and researchers classify images using labels written in natural language, without collecting training examples for each task. It runs locally through PyTorch on a CPU or CUDA GPU. The code and model weights use the MIT license.
22.2KUpdated 1 week agoMIT
#Hugging Face integration#Multilingual
Jina Embeddings is a family of models that converts content into vectors for retrieval, similarity matching, classification and clustering. It includes multilingual text models and multimodal variants for searching across different media.
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
GTE (General Text Embedding) is Alibaba’s family of downloadable models for representing text as vectors. Developers use these representations to compare queries with documents, cluster related text or supply retrieval components for larger applications.
E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.