
Jina Embeddings is a family of models that converts content into vectors for retrieval, similarity matching, classification and clustering. It includes multilingual text models and multimodal variants for searching across different media.
Downloadable weights can run locally through Transformers or SentenceTransformers. The v5-text models offer task-specific LoRA adapters, adjustable embedding dimensions and support for long inputs. Their model cards provide CPU and GPU examples; requirements vary by model size and runtime. GGUF and MLX releases provide additional options for local execution.
The v5-omni variants encode text, images, audio, video and PDFs in a shared vector space. The matching v5-text and v5-omni variants preserve compatible text embeddings, allowing an existing text index to incorporate other media. This does not make embeddings from unrelated model families interchangeable.
The v3 and v5-text-small weights are licensed under CC BY-NC 4.0. Commercial use requires the appropriate commercial terms; check the specific model’s license before deploying it. Jina also offers a hosted API and commercial On-Prem containers. The documented On-Prem offering runs offline on your own infrastructure without external calls or a license server.
The API and commercial containers support OpenAI-compatible embedding schemas. A vector store or retrieval application supplies indexing and search around the models; embeddings alone do not generate answers or source citations.
Claim this page with an email at jina.ai. Jina Embeddings gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Jina Embeddings?Promote it
Something wrong or outdated on this page?
2KUpdated 1 year ago
Docker#Batch processing#Hugging Face integration#Multilingual
Qwen3-Embedding is a family of text embedding models for developers building search and document analysis on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its multilingual support covers languages including English, Chinese, Arabic and Ukrainian, as well as programming languages.
1.9KUpdated 11 months ago
Docker#Batch processing#Hugging Face integration#ONNX
Nomic Embed Text v1.5 is an English text embedding model for developers building semantic search, document retrieval, and RAG applications on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its main distinction is adjustable embedding size: you can use smaller vectors when storage matters, with a tradeoff in retrieval quality.
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
5.8KUpdated 2 days agoApache-2.0
#Hugging Face integration#Multilingual#Quantization
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
EmbeddingGemma is a text embedding model for developers building search and document features that run on phones, laptops or tablets. Based on Gemma 3, it converts text into numerical representations so applications can find related passages by meaning. Embeddings stay on your hardware, and the model works without an internet connection.
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.