Favicon of Jina Embeddings

Jina Embeddings

Jina’s embedding models encode multilingual text and media for retrieval, with local weights, noncommercial licenses and commercial deployment options.

Screenshot of Jina Embeddings website

Jina Embeddings is a family of models that converts content into vectors for retrieval, similarity matching, classification and clustering. It includes multilingual text models and multimodal variants for searching across different media.

Downloadable weights can run locally through Transformers or SentenceTransformers. The v5-text models offer task-specific LoRA adapters, adjustable embedding dimensions and support for long inputs. Their model cards provide CPU and GPU examples; requirements vary by model size and runtime. GGUF and MLX releases provide additional options for local execution.

The v5-omni variants encode text, images, audio, video and PDFs in a shared vector space. The matching v5-text and v5-omni variants preserve compatible text embeddings, allowing an existing text index to incorporate other media. This does not make embeddings from unrelated model families interchangeable.

The v3 and v5-text-small weights are licensed under CC BY-NC 4.0. Commercial use requires the appropriate commercial terms; check the specific model’s license before deploying it. Jina also offers a hosted API and commercial On-Prem containers. The documented On-Prem offering runs offline on your own infrastructure without external calls or a license server.

The API and commercial containers support OpenAI-compatible embedding schemas. A vector store or retrieval application supplies indexing and search around the models; embeddings alone do not generate answers or source citations.

Similar to Jina Embeddings