Favicon of Qwen3-Embedding

Qwen3-Embedding

Text embedding models for self-hosted multilingual and code search, with custom vector sizes and support for CPU or NVIDIA GPU inference.

Screenshot of Qwen3-Embedding website

Qwen3-Embedding is a family of text embedding models for developers building search and document analysis on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its multilingual support covers languages including English, Chinese, Arabic and Ukrainian, as well as programming languages.

The models support document and code retrieval, text classification, and grouping related texts. They can also match passages across languages, useful for searching collections where queries and documents aren't written in the same language. The family builds on Qwen3's dense base models and includes Qwen3-Embedding-0.6B, Qwen3-Embedding-4B and Qwen3-Embedding-8B.

Developers can choose smaller output vectors to reduce storage needs, rather than keeping a fixed embedding size. Task instructions let the models adapt their embeddings to a particular retrieval problem or language. Long-text support helps with larger documents, and the companion Qwen3-Reranker models can reorder retrieved results for relevance.

Local inference works through Sentence Transformers, Transformers and vLLM. Hugging Face Text Embeddings Inference supports self-hosted Docker deployments on CPUs or NVIDIA GPUs, so embedding requests can run on infrastructure you control. The weights use Safetensors. Hosted inference is also available through Scaleway on Hugging Face; that route processes requests on a provider's infrastructure. The downloadable embedding weights use Apache-2.0. Hardware requirements grow with the selected model size; the 8B release has substantially greater memory needs than the 0.6B model.

Similar to Qwen3-Embedding