
Qwen3-Embedding is a family of text embedding models for developers building search and document analysis on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its multilingual support covers languages including English, Chinese, Arabic and Ukrainian, as well as programming languages.
The models support document and code retrieval, text classification, and grouping related texts. They can also match passages across languages, useful for searching collections where queries and documents aren't written in the same language. The family builds on Qwen3's dense base models and includes Qwen3-Embedding-0.6B, Qwen3-Embedding-4B and Qwen3-Embedding-8B.
Developers can choose smaller output vectors to reduce storage needs, rather than keeping a fixed embedding size. Task instructions let the models adapt their embeddings to a particular retrieval problem or language. Long-text support helps with larger documents, and the companion Qwen3-Reranker models can reorder retrieved results for relevance.
Local inference works through Sentence Transformers, Transformers and vLLM. Hugging Face Text Embeddings Inference supports self-hosted Docker deployments on CPUs or NVIDIA GPUs, so embedding requests can run on infrastructure you control. The weights use Safetensors. Hosted inference is also available through Scaleway on Hugging Face; that route processes requests on a provider's infrastructure. The downloadable embedding weights use Apache-2.0. Hardware requirements grow with the selected model size; the 8B release has substantially greater memory needs than the 0.6B model.
Claim this page and we'll verify you by hand. Qwen3-Embedding gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Qwen3-Embedding?Promote it
Something wrong or outdated on this page?
1.9KUpdated 11 months ago
Docker#Batch processing#Hugging Face integration#ONNX
Nomic Embed Text v1.5 is an English text embedding model for developers building semantic search, document retrieval, and RAG applications on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its main distinction is adjustable embedding size: you can use smaller vectors when storage matters, with a tradeoff in retrieval quality.
jina.aiEmbedding and Reranker Models
Docker#GGUF#LoRA#MLX
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
huggingface.coEmbedding and Reranker Models
#Batch processing#Hugging Face integration#Multilingual
22.2KUpdated 1 week agoMIT
#Hugging Face integration#Multilingual
5.8KUpdated 2 days agoApache-2.0
#Hugging Face integration#Multilingual#Quantization
Jina Embeddings is a family of models that converts content into vectors for retrieval, similarity matching, classification and clustering. It includes multilingual text models and multimodal variants for searching across different media.
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
GTE (General Text Embedding) is Alibaba’s family of downloadable models for representing text as vectors. Developers use these representations to compare queries with documents, cluster related text or supply retrieval components for larger applications.
E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.
EmbeddingGemma is a text embedding model for developers building search and document features that run on phones, laptops or tablets. Based on Gemma 3, it converts text into numerical representations so applications can find related passages by meaning. Embeddings stay on your hardware, and the model works without an internet connection.