127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
26.5KUpdated 3 weeks agoApache-2.0
macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend
MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
3.5KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multimodal input
PaliGemma is a family of downloadable vision-language models for developers and researchers building applications that work with images and text. It combines SigLIP's image processing with Gemma's language capabilities to answer questions about visual content. Its main appeal is task-specific fine-tuning: you can adapt a base model to your own image data and intended use.
1.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
PixArt-Sigma is a text-to-image diffusion model that supports generation at 2K and 4K resolutions on your own machine or server. It's aimed at developers and researchers who want pretrained models they can run themselves, along with code for training and adapting them. The Python code is open source under Apache 2.0.
25KUpdated 2 years agoApache-2.0
macOS · Web#LoRA#Multimodal input#Quantization
LLaVA is a family of vision-language models for researchers and developers who want to ask questions about images on their own hardware. It pairs a CLIP vision encoder with a language model to support image descriptions, visual reasoning and reading text in pictures. Its Python code is open source under Apache 2.0; the project places research-use restrictions on its data and checkpoints, with additional terms from the underlying models.
820Updated 1 year ago
#Quantization#Tool calling
Hunyuan-A13B is Tencent's downloadable language model for developers and researchers building self-hosted AI applications. It supports general text tasks, reasoning and agent workloads, with a choice between quick responses and more deliberate reasoning. It's part of the broader Hunyuan model family, which Tencent also offers through its website.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
2.1KUpdated 3 years agoApache-2.0
#Hugging Face integration#LoRA#Quantization
StarCoder2 is a family of code generation models for developers who want to run code completion on their own hardware or adapt a model to their code. It predicts code continuations rather than following conversational instructions, so it's suited to completion workflows rather than a chat-based coding assistant.
7KUpdated 11 months agoMIT
#Hugging Face integration#Multilingual
DeepSeek-Coder is a family of code language models for developers who want downloadable models for their own coding tools. It generates code, completes existing code, and fills gaps within a file using the surrounding context. DeepSeek also offers a hosted chat interface; the downloadable weights are separate from that service.
3.9KUpdated 1 month agoMIT
#Hugging Face integration#Ollama integration#Works offline
Phi-4 is Microsoft's small language model for complex reasoning and math problem solving. It's for developers building generative AI applications who want to run a model on their own hardware, rather than depend on a cloud connection. The original Phi-4 has 14 billion parameters and uses the MIT license; hardware needs depend on the chosen model and precision.
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
7.8KUpdated 2 years agoApache-2.0
Docker#Hugging Face integration#llama.cpp backend#Multilingual
Yi is a family of open-weight language models from 01.AI for developers, researchers, and businesses that want to run English and Chinese models on their own hardware. It includes models for conversation and base models for fine-tuning, with license terms that depend on the release. Yi-1.5 code and weights use Apache 2.0; the original Yi releases have separate community terms.
3.2KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.
7.3KUpdated 1 year agoApache-2.0
#Hugging Face integration#Tool calling#Web search
InternLM is a family of downloadable language models for developers and researchers building their own AI applications. It includes models for general conversation, complex reasoning and coding, with separate base and chat variants for customization or use in an assistant.
9.2KUpdated 2 weeks agoApache-2.0
#ControlNet#LoRA#Multimodal input
Sana is an open-source framework for running image and video generation on your own hardware, with image models small enough for laptop GPUs. It's aimed at creators who want local AI generation and developers who need training and inference pipelines for their own models. The code uses the Apache 2.0 license.
2.5KUpdated 1 year agoMIT
Web#Hugging Face integration
HiDream-I1 is an open-source text-to-image model for people who want to generate images on their own hardware or build image generation into a Python application. It uses MIT licensing and supports local inference through CUDA, making it an option for developers and creators with NVIDIA GPU hardware.
34.4KUpdated 6 months agoMIT
#Batch processing#Multimodal input
CLIP is an open-source image and text model that lets developers and researchers classify images using labels written in natural language, without collecting training examples for each task. It runs locally through PyTorch on a CPU or CUDA GPU. The code and model weights use the MIT license.
1.3KUpdated 7 months agoApache-2.0
Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API
JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.
5.6KUpdated 2 years agoApache-2.0
macOS#Hugging Face integration
Parler-TTS generates speech locally, with text descriptions that control how the voice sounds. It's a Python library for developers building speech into applications and researchers who want to train or adapt a TTS model. The library uses the Apache 2.0 license and can run on CPU or CUDA GPUs, with support for Apple Silicon.
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
23.3KUpdated 1 year agoApache-2.0
macOS · Windows · Web#Batch processing#Image-to-image#Inpainting
IOPaint is a free, self-hosted AI image editor for people who want to remove objects, replace parts of a picture or extend it beyond its original edges on their own hardware. The project is archived and no longer maintained. It's open source under Apache 2.0, with a browser interface and support for CPU, GPU and Apple Silicon hardware.
14.2KUpdated 5 days ago
#Hugging Face integration#Multimodal input
OpenCLIP is a Python and PyTorch library for developers and researchers who want to match images with text on their own hardware. It implements OpenAI's CLIP approach: images and descriptions become numerical representations that the model can compare. This supports image search and zero-shot classification, where text labels define the categories without a separate classifier trained for each task.
14.9KUpdated 2 years agoApache-2.0
macOS · Windows · Docker#Streaming inference#Voice cloning
Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.
15.4KUpdated 6 days agoMIT
Web#Quantization
ggml is an open-source library for developers building software that runs machine learning models on their own hardware. It provides tensor computation in C/C++ without dependencies, with backends for CPUs, GPUs, NPUs and browsers. The project also documents the GGUF model file format.
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
4.6KUpdated 7 months agoMIT
Windows · Linux#Batch processing#Quantization#Speculative decoding
ExLlamaV2 is a local LLM inference library for developers and people hosting models on their own consumer GPUs. ExLlamaV2 is archived and no longer maintained; development continues in ExLlamaV3. The V2 library is free and open source under the MIT license, runs on Windows and Linux, and uses NVIDIA GPUs through CUDA. It supports multiple GPUs.
11.2KUpdated 5 months agoMIT
VS Code#Code execution#LLM tracing#Visual workflows
Prompt flow is an MIT-licensed, open-source toolkit for developers who build LLM applications and need to test their behavior before deployment. Its development tools run locally, while an optional cloud version in Azure AI supports team collaboration. Feature development has ended.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
17.8KUpdated 2 years agoMIT
Web#Batch processing#Hugging Face integration#Multimodal input
Janus-Pro is a multimodal AI model from DeepSeek that answers questions about images and creates pictures from text prompts. It runs on your own hardware and suits developers and researchers who want both capabilities in one model. A local Gradio demo provides a browser interface, while Hugging Face hosts a separate online demo.
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
39.3KUpdated 2 years agoMIT
#Hugging Face integration#Multilingual
Bark is Suno's local text-to-audio model for developers and researchers who want to generate speech alongside other sounds. It can produce laughter, crying, music and background noise within its output. Its generative approach suits audio experiments, though it can depart from the supplied script and doesn't guarantee clean, studio-quality speech.
1.1KUpdated 7 months agoAGPL-3.0
Web · Browser Extension#Multilingual#Multimodal input#Ollama integration
NativeMind brings a local AI assistant into Chrome for people who want help with webpages, documents, and writing without sending that content to a cloud model. The browser extension connects to Ollama on your machine, keeping prompts and AI processing on-device. It requires no account and uses the AGPL-3.0 open-source license.
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
2KUpdated 7 months agoMIT
macOS · Windows · Linux · Web#MCP#RAG
NotebookLlama is a self-hosted NotebookLM alternative for people who want to work with documents through an app they can host and modify. Its browser interface runs locally, while LlamaCloud provides document extraction and indexing. It requires cloud services.
3.5KUpdated 22 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
960Updated 7 months agoApache-2.0
#Hugging Face integration#LoRA#Quantization
HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.
privatellm.appAutomation and No-Code AI
macOS · iOS#Multilingual#Quantization#Works offline
Private LLM runs AI chat entirely on your iPhone, iPad, or Mac. It's an app for people who want to use language models without sending their conversations to a cloud service. After the first model download, it works offline and requires no account. Conversations stay on-device, with no tracking or logs.
31.4KUpdated 1 week agoApache-2.0
macOS#Distributed execution#ONNX
PyTorch Lightning is a Python framework for researchers and developers who want to pretrain or fine-tune models on their own hardware or in the cloud. It handles repetitive training code while leaving model logic under your control. The framework is open source under Apache 2.0.
5.7KUpdated 2 weeks agoMIT
macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows
GLaDOS is a local AI voice assistant modeled on the sarcastic character from Valve's Portal games. It's for people who want a conversational companion on their own hardware, with camera awareness and connections to home automation. The Python project is open source under the MIT license and runs on Linux and Windows. macOS support is experimental.
2.8KUpdated 1 week agoApache-2.0
Linux#Distributed execution#Hugging Face integration
Nanotron is a Python library for researchers and developers who want to pretrain language models on their own datasets and GPU infrastructure. Built on PyTorch, it supports NVIDIA CUDA GPUs and training across multiple servers with Slurm. It's open source under Apache 2.0.
1.2KUpdated 12 months agoMIT
macOS · Windows · Linux · Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API
Hollama is an open-source LLM chat app whose interface runs entirely in your browser. It's for people who want to chat with local AI through Ollama or connect to OpenAI servers, with support for multiple server connections. The interface stores data locally in the browser; the connected server handles model requests, so where inference runs depends on the server you choose.