407Updated 2 days agoMIT
macOS#MLX#Multimodal input#OpenAI-compatible API
Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.
13.2KUpdated 2 days agoApache-2.0
#ControlNet#Image-to-image#Inpainting
DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.
34.6KUpdated 1 day agoApache-2.0
macOS#ControlNet#Hugging Face integration#Image-to-image
Diffusers is an open-source Python library for developers and researchers who want to run diffusion models on their own hardware or build generation features into an application. It uses PyTorch and supports image, video and audio generation. The library is licensed under Apache 2.0 and supports Apple Silicon.
130KUpdated 39 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
28.6KUpdated 24 hours agoMIT
macOS · Linux#Distributed execution#LoRA
MLX is a machine learning array framework for researchers and developers building models on their own hardware. Its distinctive feature on Apple silicon is shared CPU and GPU memory: both processors can work on the same arrays without copying data between them. It's open source under the MIT license.
54KUpdated 2 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
166.8KUpdated 1 day agoApache-2.0
#Hugging Face integration#Multimodal input
Transformers is a Python library for developers and researchers who want to run pretrained AI models or train their own on hardware they control. It covers language, images, audio, video and multimodal work through a shared way of defining models. The library runs in a local Python environment; pretrained checkpoints are available from the separate Hugging Face Hub.
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
3.9KUpdated 2 years agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#Multimodal input
Riffusion is a Python library for generating music and audio on your own hardware using Stable Diffusion. It's for developers and musicians who want to experiment with text-driven sound generation or build it into an app. The hobby project is no longer actively maintained.
550Updated 2 years agoMPL-2.0
Web#Multilingual
Bergamot Translator is a C++ library for running neural machine translation on your own device. Developers can integrate it into native applications or build it for WebAssembly. Separate browser integrations, such as the TranslateLocally extension, make the engine accessible to end users. The software is free and open source under the Mozilla Public License 2.0.
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
2.3KUpdated 1 year agoMIT
Linux#Batch processing#GGUF#Hugging Face integration
AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
16.2KUpdated 1 day agoApache-2.0
Windows · iOS · Android#Image-to-image#Multimodal input#ONNX
MNN is a lightweight C++ framework for developers who want AI models to run on phones, PCs and embedded devices. It handles inference and training on the device, with a focus on small application footprints and hardware acceleration. The project is open source under Apache 2.0, and Alibaba uses it in apps including Taobao, Youku and DingTalk.
5.6KUpdated 2 years agoApache-2.0
macOS#Hugging Face integration
Parler-TTS generates speech locally, with text descriptions that control how the voice sounds. It's a Python library for developers building speech into applications and researchers who want to train or adapt a TTS model. The library uses the Apache 2.0 license and can run on CPU or CUDA GPUs, with support for Apple Silicon.
14.2KUpdated 5 days ago
#Hugging Face integration#Multimodal input
OpenCLIP is a Python and PyTorch library for developers and researchers who want to match images with text on their own hardware. It implements OpenAI's CLIP approach: images and descriptions become numerical representations that the model can compare. This supports image search and zero-shot classification, where text labels define the categories without a separate classifier trained for each task.
15.4KUpdated 6 days agoMIT
Web#Quantization
ggml is an open-source library for developers building software that runs machine learning models on their own hardware. It provides tensor computation in C/C++ without dependencies, with backends for CPUs, GPUs, NPUs and browsers. The project also documents the GGUF model file format.
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
4.6KUpdated 7 months agoMIT
Windows · Linux#Batch processing#Quantization#Speculative decoding
ExLlamaV2 is a local LLM inference library for developers and people hosting models on their own consumer GPUs. ExLlamaV2 is archived and no longer maintained; development continues in ExLlamaV3. The V2 library is free and open source under the MIT license, runs on Windows and Linux, and uses NVIDIA GPUs through CUDA. It supports multiple GPUs.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
960Updated 7 months agoApache-2.0
#Hugging Face integration#LoRA#Quantization
HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.
2.8KUpdated 9 months agoApache-2.0
Windows · Linux#Batch processing#ONNX#Voice activity detection
openWakeWord is a Python library for developers building voice interfaces that listen locally for a chosen word or phrase. It includes English models for triggers such as "hey jarvis" and "alexa", plus phrases for weather and timers. The code uses Apache 2.0. Included pretrained models use CC-BY-NC-SA-4.0, which restricts commercial use.
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
41.4KUpdated 3 days agoApache-2.0
#Distributed execution
Colossal-AI is a Python framework for developers and researchers training or serving large AI models on their own GPU hardware. It addresses the memory and computing demands of models that are difficult to fit on a single GPU, with tools for distributing work across a cluster. It's open source under Apache 2.0.
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
23.7KUpdated 2 years agoMIT
#Multimodal input
MusicGen is Meta AI's music generation model within AudioCraft, a PyTorch library for developers and audio researchers who want to generate music in their own computing environment. It creates music from text descriptions and can use a melody to guide the result. AudioCraft includes both inference and training code, so it's suited to people building audio tools or studying music generation.
7.7KUpdated 2 years agoMIT
Docker · Web#Multilingual
MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.
37.1KUpdated 19 hours agoApache-2.0
iOS · Android · Web
MediaPipe is an open-source toolkit for developers adding on-device AI to applications on Android, iOS, the web, desktop and edge devices. It pairs pretrained models with APIs for specific tasks, so developers can use existing solutions or customize them for their applications. The project uses the Apache 2.0 license.
30KUpdated 10 months agoApache-2.0
Windows#Multilingual
EasyOCR is a Python OCR library for developers who want to extract text from images on their own hardware. It reads text in photographs and dense documents, so it can serve both scene-text recognition and document processing. It's open source under Apache 2.0.
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.
34.7KUpdated 2 days agoApache-2.0
Detectron2 is an open-source Python library for developers and researchers building computer vision applications. It provides algorithms for locating objects in images and segmenting image regions, with support for training models and building research projects on top of the library. Facebook AI Research developed it as the successor to Detectron and maskrcnn-benchmark.