2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
jina.aiEmbedding and Reranker Models
Docker#GGUF#LoRA#MLX
Jina Embeddings is a family of models that converts content into vectors for retrieval, similarity matching, classification and clustering. It includes multilingual text models and multimodal variants for searching across different media.
8.4KUpdated 4 months agoApache-2.0
#Batch processing#Hugging Face integration#Multilingual
Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.
22.2KUpdated 1 week agoMIT
#Hugging Face integration#Multilingual
E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.
14KUpdated 5 months ago
macOS · Windows · Linux · Web#GGUF#LM Studio integration#MCP
Open-LLM-VTuber is a local AI companion for people who want a character they can talk to, with a Live2D avatar that responds through speech and expressions. It runs on Windows, macOS and Linux through web and desktop clients. With local models for speech and language processing, it works fully offline and keeps conversations on your device. Cloud APIs are optional alternatives that send the corresponding processing to external services.
4.7KUpdated 1 year agoMIT
#Hugging Face integration#Multilingual#Voice cloning
WhisperSpeech is a local text-to-speech system that uses OpenAI Whisper as the basis for generating speech. It's for developers and speech researchers who want to work with downloadable models on their own hardware. It supports voice cloning.
7.2KUpdated 2 years agoApache-2.0
macOS · Linux · Docker · Web#Multilingual#Voice cloning
Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
4.6KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Multilingual
AutoTrain trains custom machine learning models from your own data through a no-code interface. It's for people who need to fine-tune an LLM or build a classifier without writing a training pipeline. The local AutoTrain Advanced project is no longer maintained, so it won't receive bug fixes or new features.
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.
6KUpdated 6 months agoMIT
Docker · Web#Multilingual#Ollama integration#OpenAI-compatible API
Paperless-AI is a self-hosted extension for Paperless-ngx users who want automatic document sorting and chat with their archive. It requires an existing Paperless-ngx instance and runs in Docker, with a browser interface for reviewing and processing documents. The project is no longer maintained.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
1.7KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#Multilingual#Persistent memory
RisuAI is an open-source AI roleplay client for people who want to create characters, build fictional worlds and chat with several characters together. It runs on Windows, macOS, Linux, Android and in a browser. You can also host the web app yourself with Docker.
5.4KUpdated 3 months ago
macOS · Windows · Linux#MCP#Multilingual#Ollama integration
5ire is a free desktop AI assistant that combines chat, a local document knowledge base and MCP tools. It runs on macOS, Windows and Linux, with Mac downloads for Apple Silicon and Intel. It's for people who want to use their own documents and external tools alongside conversations with local or cloud models.
1.1KUpdated 2 years agoApache-2.0
Web#Batch processing#Hugging Face integration#LoRA
CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.
huggingface.coOpen-Weight LLMs
#Hugging Face integration#LoRA#Multilingual
Jamba is AI21's language model family for teams building AI applications on their own servers. The documented Large 1.7 model combines Mamba state-space models with Transformer attention to process long context efficiently. A 256K-token context window makes it relevant for work that depends on lengthy documents, such as investment research, due diligence and reviewing procurement responses.
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
15.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
26.5KUpdated 3 weeks agoApache-2.0
macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend
MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
7KUpdated 11 months agoMIT
#Hugging Face integration#Multilingual
DeepSeek-Coder is a family of code language models for developers who want downloadable models for their own coding tools. It generates code, completes existing code, and fills gaps within a file using the surrounding context. DeepSeek also offers a hosted chat interface; the downloadable weights are separate from that service.
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
7.8KUpdated 2 years agoApache-2.0
Docker#Hugging Face integration#llama.cpp backend#Multilingual
Yi is a family of open-weight language models from 01.AI for developers, researchers, and businesses that want to run English and Chinese models on their own hardware. It includes models for conversation and base models for fine-tuning, with license terms that depend on the release. Yi-1.5 code and weights use Apache 2.0; the original Yi releases have separate community terms.
3.2KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
39.3KUpdated 2 years agoMIT
#Hugging Face integration#Multilingual
Bark is Suno's local text-to-audio model for developers and researchers who want to generate speech alongside other sounds. It can produce laughter, crying, music and background noise within its output. Its generative approach suits audio experiments, though it can depart from the supplied script and doesn't guarantee clean, studio-quality speech.
1.1KUpdated 7 months agoAGPL-3.0
Web · Browser Extension#Multilingual#Multimodal input#Ollama integration
NativeMind brings a local AI assistant into Chrome for people who want help with webpages, documents, and writing without sending that content to a cloud model. The browser extension connects to Ollama on your machine, keeping prompts and AI processing on-device. It requires no account and uses the AGPL-3.0 open-source license.
privatellm.appAutomation and No-Code AI
macOS · iOS#Multilingual#Quantization#Works offline
Private LLM runs AI chat entirely on your iPhone, iPad, or Mac. It's an app for people who want to use language models without sending their conversations to a cloud service. After the first model download, it works offline and requires no account. Conversations stay on-device, with no tracking or logs.
13.2KUpdated 1 year ago
Linux · Docker#Multilingual#Multimodal input
Wav2Lip is a local AI lip-sync tool that changes a face's mouth movements in an existing video to match supplied speech. It's for researchers and people making academic or personal video projects who want to process their own files. Commercial use is prohibited under the project's stated terms because its models were trained on the LRS2 dataset.
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
3.1KUpdated 1 year agoAGPL-3.0
Web#Multilingual#Works offline
Whishper is a self-hosted speech-to-text app for people who need transcripts or translated subtitles from audio and video. Its browser interface brings transcription, translation and subtitle editing together, with all three running on your own machine. It can work offline, so local media doesn't need to go to a cloud transcription service.