2.7KUpdated 2 weeks agoMIT
Android#GGUF#Hugging Face integration#llama.cpp backend
Maid is an Android AI chat app for people who want to run models on their phone and access remote models in the same app. It runs GGUF models locally through llama.cpp without an internet connection. It's open source under the MIT license, with no ads or telemetry.
16.8KUpdated 1 day agoMIT
Docker · Web#Hugging Face integration#Multi-user access#ONNX
CVAT is a browser-based data annotation platform for teams building computer vision datasets. Its open-source Community edition runs on your own infrastructure with Docker and uses the MIT license. CVAT Online is hosted by CVAT, while the Enterprise offering runs in an organization's own cloud or internal environment.
3KUpdated 5 days ago
Linux · iOS#Hugging Face integration#LoRA#Quantization
TorchAO is a PyTorch library for developers who want to train or run models on their own hardware with less memory and faster computation. It reduces the precision of model weights and activations, with options for language models and image or video generation. Its PyTorch integration works with torch.compile and FSDP2 across most Hugging Face PyTorch models.
26.1KUpdated 5 hours ago
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llamafile puts an LLM and the software that runs it into a single executable. It’s for people who want to run models locally or share them with others without asking each recipient to set up a separate runtime. The project is open source under Apache 2.0.
12.6KUpdated 3 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#ControlNet#Hugging Face integration#LoRA
Kohya's GUI lets you train and fine-tune image generation models on your own GPU-equipped computer through a browser interface. It's for artists and model makers who want to teach a model a particular style or subject while controlling the training settings. The interface builds on Kohya's Stable Diffusion training scripts, with a command-line interface available too.
8.9KUpdated 2 weeks agoAGPL-3.0
macOS · Windows · Linux#Hugging Face integration#LoRA
Stability Matrix is an open source desktop app for people who use more than one Stable Diffusion interface. It manages local installations of ComfyUI, Automatic1111, Fooocus, Forge, and InvokeAI, so you can try different workflows without maintaining a separate model collection for each one. It runs on Windows, macOS, and Linux under the AGPL-3.0 license.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
1.5KUpdated 2 years agoMIT
#ControlNet#Hugging Face integration#Image-to-image
Stable Diffusion 3.5 is a family of text-to-image models for people building image tools or producing visual work on their own infrastructure. It generates photography, paintings, line art and 3D-style images from prompts, with an emphasis on following the requested subject and composition.
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
937Updated 2 years agoApache-2.0
#Hugging Face integration#Multimodal input#Works offline
Molmo is Ai2's family of vision-language models, with code for running and training models on your own hardware. It's for developers and researchers who need to work with images and text, adapt a model, or evaluate it against visual tasks. The Python codebase is open source under Apache 2.0 and builds on OLMo, adding image encoding and generative evaluation.
24.3KUpdated 4 days agoBSD-2-Clause
macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual
WhisperX is an open source speech-to-text tool for people transcribing interviews, meetings, and long recordings on their own computer. It builds on OpenAI's Whisper to produce transcripts with word-level timestamps and optional speaker labels.
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.
390Updated 1 day agoMIT
Docker#Home Assistant integration#Hugging Face integration#Multilingual
Wyoming Faster Whisper is a local speech-to-text server for Home Assistant and other clients that use the Wyoming protocol. It turns spoken audio into text on your own hardware, with support for names specific to your home. It's open source under the MIT license and runs as a Home Assistant add-on, a Docker container, or a local Python service.
23.7KUpdated 1 day agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
verl is a Python library for teams training large language models on their own GPU infrastructure. It's the open-source implementation of HybridFlow, aimed at researchers and engineers who need reinforcement learning after initial model training. It uses the Apache 2.0 license.
8.5KUpdated 2 days agoMIT
iOS · Android#GGUF#Hugging Face integration#llama.cpp backend
PocketPal AI is an open source assistant for people who want to run language models on a phone or tablet. It works on iOS, iPadOS and Android. Once you've downloaded a model, you can chat offline without an account, and your prompts, replies and documents stay on your device. The app is licensed under MIT.
21.8KUpdated 1 week agoMIT
macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization
Buzz transcribes and translates speech on your own computer using OpenAI's Whisper. It's for people who need transcripts or subtitles from recordings, plus live captions from a microphone. Local transcription works offline; the optional OpenAI Whisper API sends audio to a cloud service.
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
10.1KUpdated 2 weeks agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.
2.9KUpdated 24 hours agoMIT
Web · VS Code#Code execution#Hugging Face integration#MCP
Inspect AI is a Python framework for researchers and developers testing language models and AI agents. Developed by the UK AI Security Institute and Meridian Labs, it evaluates coding, reasoning, knowledge, behavior and multimodal understanding, including tasks where agents must take actions to succeed.
3.9KUpdated 1 day agoApache-2.0
#Hugging Face integration
Hugging Face CLI is a terminal client for finding, downloading, and sharing models and datasets on the Hugging Face Hub. It's for developers working with local AI, model authors publishing their work, and coding agents that need access to Hub resources. The client runs on your machine and connects to Hugging Face's hosted platform.
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
19.1KUpdated 4 months ago
macOS · Windows · Linux · Web#Hugging Face integration
LivePortrait animates a still portrait using facial expressions and head movements from a driving video. It's for creators who want to animate faces or edit motion in portrait videos on their own hardware. It works with realistic photos as well as portraits in oil paintings, sculptures and 3D renders.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
37.3KUpdated 1 day agoAGPL-3.0
macOS · Windows · Docker · Web#Batch processing#Hugging Face integration#MCP
PDFMathTranslate translates scientific PDFs while keeping their page layout, formulas, charts, contents pages and annotations. It's for researchers, students and others who need to read papers in another language without losing the relationship between the text and its figures. It produces both translated PDFs and bilingual documents for comparison with the original.
9.4KUpdated 2 weeks agoApache-2.0
#AI red teaming#GGUF#Hugging Face integration
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.
13.7KUpdated 11 months agoMIT
Linux · Web#Hugging Face integration#Multimodal input
TRELLIS is a local AI model for generating 3D assets from images or text prompts, aimed at 3D artists and researchers exploring asset creation. It can produce meshes, radiance fields and 3D Gaussians from the same underlying representation, so you can choose an output suited to your rendering or editing work.
5.8KUpdated 5 months agoBSD-3-Clause
#Hugging Face integration#LoRA#Quantization
torchtune is a Python library for developers and researchers who want to adapt LLMs on their own GPU hardware using PyTorch. Its editable training recipes suit work that needs control over the training code and model implementations. The project is no longer actively maintained.
5.1KUpdated 1 week agoApache-2.0
macOS · Linux · Docker#Batch processing#Hugging Face integration#LLM tracing
Text Embeddings Inference is a self-hosted server for developers who need text embeddings for search and retrieval applications. It serves models through a REST API on your own hardware and can run offline once model weights are downloaded. The Rust project is open source under Apache 2.0.
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
5.8KUpdated 3 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.