9.2KUpdated 7 days agoApache-2.0
Docker#Image-to-image#Inpainting#Multimodal input
ModelScope combines a hosted model and dataset hub with a Python library you can run locally. It's for developers and researchers who want to use AI models in their own applications, fine-tune them on their own data, or compare their performance. The library is open source under Apache 2.0.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
1.8KUpdated 2 months agoApache-2.0
macOS#MLX#Streaming inference
Magenta RealTime 2 is a local AI music model and synthesis engine for musicians and developers who want to play or build AI musical instruments on a laptop. It generates streaming audio in real time, with open weights and code under the Apache 2.0 license.
29.9KUpdated 3 weeks ago
macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization
InsightFace is a face analysis toolkit for developers and teams building identity verification, access control, or face editing software. The code uses the MIT license. Its Python tools and self-hosted recognition server run inference on your own hardware. It also offers commercial models and API access for face swapping and deepfake detection.
9.8KUpdated 5 months agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration
PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.
10.6KUpdated 2 years agoMIT
macOS · Windows · Linux · Docker#Hugging Face integration
Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.
23.5KUpdated 23 hours agoMIT
Docker
DeepFace is a Python library for developers building face recognition into their own applications. It runs locally or as a self-hosted API, including through Docker. A separate managed API at deepface.dev handles processing on hosted infrastructure. The library is open source under the MIT license.
1.6KUpdated 9 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
33KUpdated 3 years agoApache-2.0
MMDetection is a Python toolkit for researchers and developers building object detection and image segmentation systems. Part of OpenMMLab, it combines ready-made model architectures with interchangeable components for custom models. It's open source under Apache 2.0.
1.7KUpdated 2 days ago
Linux#Multimodal input#Quantization
RKLLM is a software stack for developers building local AI applications on Rockchip hardware. It uses the chip's neural processing unit (NPU) to run language and multimodal models on development boards, with support for the RK3588, RK3576, RK3562 and RV1126B series.
3KUpdated 5 days ago
Linux · iOS#Hugging Face integration#LoRA#Quantization
TorchAO is a PyTorch library for developers who want to train or run models on their own hardware with less memory and faster computation. It reduces the precision of model weights and activations, with options for language models and image or video generation. Its PyTorch integration works with torch.compile and FSDP2 across most Hugging Face PyTorch models.
43.2KUpdated 1 day agoApache-2.0
Windows#Distributed execution
DeepSpeed is an open-source library for developers and researchers training or running large AI models on their own hardware or compute clusters. It works with PyTorch and focuses on memory use, training speed, and distributing work across GPUs. It's licensed under Apache 2.0.
22KUpdated 27 minutes agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
9.1KUpdated 1 year agoApache-2.0
macOS · Windows#Batch processing#Multilingual#ONNX
Kokoro is a text-to-speech model and inference library for developers who want to generate speech on their own hardware or servers. Its compact Kokoro-82M model suits personal projects and production applications, with Apache 2.0 licensing for both the library and model weights.
4.7KUpdated 5 days agoMIT
#Batch processing#Multilingual#Quantization
CTranslate2 is an open-source C++ and Python library for developers running Transformer models on their own hardware or servers. It handles translation, text generation, text encoding and speech recognition. Its custom runtime focuses on reducing inference time and memory use compared with general-purpose deep learning frameworks.
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
23.2KUpdated 24 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
3.9KUpdated 4 months agoMIT
Web#Hugging Face integration
Stable Audio Tools is an MIT-licensed Python toolkit for developers and audio researchers who want to generate audio on their own hardware or train models on their own recordings. It combines model inference with training and fine-tuning, so you can work with pretrained models or build a model around a specific audio dataset.
8.1KUpdated 3 days agoApache-2.0
#Batch processing#Distributed execution#Hugging Face integration
LMDeploy is an open-source toolkit for developers serving language and vision-language models on their own hardware. It combines model compression with inference and self-hosted APIs, so teams can use it for batch processing or as the model backend for an application. It uses the Apache 2.0 license.
14.7KUpdated 21 hours ago
Docker#Batch processing#Distributed execution#LoRA
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
8.5KUpdated 4 weeks agoMIT
macOS · Windows · Linux#LoRA#Quantization
bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.
19.5KUpdated 1 week agoApache-2.0
Docker#LoRA#Multimodal input#Prompt caching
KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.
2.2KUpdated 3 days agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Guardrails
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.
10.6KUpdated 3 months agoMIT
#Hugging Face integration#Speaker diarization#Voice activity detection
pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
10.9KUpdated 24 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.
3.2KUpdated 24 hours agoApache-2.0
#Batch processing#Multilingual#ONNX
FastEmbed is a Python library that generates embeddings on your own hardware for semantic search and retrieval-augmented generation (RAG). It's for developers who need to turn text into searchable vectors without relying on a cloud embedding API. It can run on a CPU or use GPU acceleration, and its Apache 2.0 license makes it open source.
9.9KUpdated 1 day agoApache-2.0
#Distributed execution
Accelerate is a Python library for developers and researchers who write their own PyTorch training loops and want to use the same code on a local machine or a distributed cluster. It handles the hardware-specific work while leaving the training logic under your control.