2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
14.1KUpdated 2 weeks agoMIT
macOS#Batch processing#GGUF#Hugging Face integration
lm-evaluation-harness lets researchers and model developers compare language models using shared academic benchmarks and public prompts. It runs evaluations against local models and benchmarks, or sends requests to a hosted model API. EleutherAI's Python framework is open source under the MIT license and powers Hugging Face's Open LLM Leaderboard.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
26.5KUpdated 3 weeks agoApache-2.0
macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend
MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
25KUpdated 2 years agoApache-2.0
macOS · Web#LoRA#Multimodal input#Quantization
LLaVA is a family of vision-language models for researchers and developers who want to ask questions about images on their own hardware. It pairs a CLIP vision encoder with a language model to support image descriptions, visual reasoning and reading text in pictures. Its Python code is open source under Apache 2.0; the project places research-use restrictions on its data and checkpoints, with additional terms from the underlying models.
820Updated 1 year ago
#Quantization#Tool calling
Hunyuan-A13B is Tencent's downloadable language model for developers and researchers building self-hosted AI applications. It supports general text tasks, reasoning and agent workloads, with a choice between quick responses and more deliberate reasoning. It's part of the broader Hunyuan model family, which Tencent also offers through its website.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
2.1KUpdated 3 years agoApache-2.0
#Hugging Face integration#LoRA#Quantization
StarCoder2 is a family of code generation models for developers who want to run code completion on their own hardware or adapt a model to their code. It predicts code continuations rather than following conversational instructions, so it's suited to completion workflows rather than a chat-based coding assistant.
7.8KUpdated 2 years agoApache-2.0
Docker#Hugging Face integration#llama.cpp backend#Multilingual
Yi is a family of open-weight language models from 01.AI for developers, researchers, and businesses that want to run English and Chinese models on their own hardware. It includes models for conversation and base models for fine-tuning, with license terms that depend on the release. Yi-1.5 code and weights use Apache 2.0; the original Yi releases have separate community terms.
9.2KUpdated 2 weeks agoApache-2.0
#ControlNet#LoRA#Multimodal input
Sana is an open-source framework for running image and video generation on your own hardware, with image models small enough for laptop GPUs. It's aimed at creators who want local AI generation and developers who need training and inference pipelines for their own models. The code uses the Apache 2.0 license.
1.3KUpdated 7 months agoApache-2.0
Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API
JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.
15.4KUpdated 6 days agoMIT
Web#Quantization
ggml is an open-source library for developers building software that runs machine learning models on their own hardware. It provides tensor computation in C/C++ without dependencies, with backends for CPUs, GPUs, NPUs and browsers. The project also documents the GGUF model file format.
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
4.6KUpdated 7 months agoMIT
Windows · Linux#Batch processing#Quantization#Speculative decoding
ExLlamaV2 is a local LLM inference library for developers and people hosting models on their own consumer GPUs. ExLlamaV2 is archived and no longer maintained; development continues in ExLlamaV3. The V2 library is free and open source under the MIT license, runs on Windows and Linux, and uses NVIDIA GPUs through CUDA. It supports multiple GPUs.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
3.5KUpdated 20 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
960Updated 7 months agoApache-2.0
#Hugging Face integration#LoRA#Quantization
HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.
privatellm.appAutomation and No-Code AI
macOS · iOS#Multilingual#Quantization#Works offline
Private LLM runs AI chat entirely on your iPhone, iPad, or Mac. It's an app for people who want to use language models without sending their conversations to a cloud service. After the first model download, it works offline and requires no account. Conversations stay on-device, with no tracking or logs.
5.1KUpdated 22 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
2.1KUpdated 3 days ago
Windows · Linux#LoRA#Quantization
Musubi Tuner is a Python toolkit for training LoRA adapters for image and video generation models on your own hardware. It's aimed at people who want to customize these models using their own datasets and are comfortable working with training scripts. It also includes image and video generation scripts for supported architectures.
752Updated 3 weeks agoApache-2.0
Linux · Docker · Web · Browser Extension#Batch processing#OpenAI-compatible API#Quantization
Meme Search is a free, self-hosted web app for people who want to find memes in their own collection by image content or text. It runs locally through Docker and uses AI descriptions to make images searchable, even when their filenames aren't useful. The code uses the Apache 2.0 license.
2KUpdated 2 days agoGPL-3.0
Windows · Linux#Distributed execution#LoRA#Quantization
diffusion-pipe is a local diffusion model training tool for people fine-tuning image and video models on their own GPU hardware. Its main distinction is that it can divide a model across several GPUs when it won't fit on one, while also distributing training work across GPUs. The Python project is open source under GPL-3.0 and uses DeepSpeed.
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
2.3KUpdated 1 year agoMIT
macOS · iOS#MLX#Quantization#Works offline
Fullmoon is an open source chat app for people who want to run language models on their Apple devices and keep conversations local. It supports iOS, iPadOS, macOS and visionOS, with on-device inference optimized for Apple silicon. You can chat fully offline, and the app saves your chat history locally.
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.
1.1KUpdated 2 years agoGPL-3.0
macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization
WhisperWriter turns microphone speech into text and types it into the window you're working in. It's for people who want voice input in their existing desktop apps, with a choice between transcription on their own computer and an external service. The Python app runs on Windows, macOS and Linux and uses the GPL-3.0 open-source license.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
214Updated 3 weeks agoMIT
Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual
Wyoming Piper connects Piper's local text-to-speech engine to Home Assistant and other clients that use the Wyoming protocol. It's for people building a voice assistant on their own hardware who need speech generation as a self-hosted service. The project is open source under the MIT license.
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
29.9KUpdated 3 weeks ago
macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization
InsightFace is a face analysis toolkit for developers and teams building identity verification, access control, or face editing software. The code uses the MIT license. Its Python tools and self-hosted recognition server run inference on your own hardware. It also offers commercial models and API access for face swapping and deepfake detection.
9.8KUpdated 5 months agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration
PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.