47.7KUpdated 1 month agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA
text-generation-webui, also called TextGen, runs language models on your own hardware through a desktop app or a self-hosted browser interface. It's for people who want private chat and writing tools, and developers who need a local model API. It works offline without telemetry; web search and page fetching use the internet.
26.1KUpdated 5 hours ago
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llamafile puts an LLM and the software that runs it into a single executable. It’s for people who want to run models locally or share them with others without asking each recipient to set up a separate runtime. The project is open source under Apache 2.0.
13KUpdated 1 year agoAGPL-3.0
Windows · Web#ControlNet#GGUF#Image-to-image
Stable Diffusion WebUI Forge runs image generation on your own hardware through a browser interface. It builds on Stable Diffusion WebUI and suits people who want its image creation tools with more control over GPU memory use, as well as developers extending those tools. It's open source under AGPL-3.0.
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
8.5KUpdated 2 days agoMIT
iOS · Android#GGUF#Hugging Face integration#llama.cpp backend
PocketPal AI is an open source assistant for people who want to run language models on a phone or tablet. It works on iOS, iPadOS and Android. Once you've downloaded a model, you can chat offline without an account, and your prompts, replies and documents stay on your device. The app is licensed under MIT.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
9.4KUpdated 2 weeks agoApache-2.0
#AI red teaming#GGUF#Hugging Face integration
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.
5.8KUpdated 2 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
5.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend
llama-swap is a self-hosted proxy for people running several AI models on their own hardware. It starts the model server a request needs and swaps out another when necessary, so you don't have to keep every model loaded or manage separate API connections in your apps.
33.5KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#GGUF#ONNX
Netron displays neural network and machine learning models as visual graphs. Developers and researchers can inspect a model's structure through desktop apps for macOS, Windows and Linux, or through the browser viewer at netron.app.
2.8KUpdated 1 week agoAGPL-3.0
Android#GGUF#llama.cpp backend#Ollama integration
ChatterUI is an Android chat app for people who want to run a local LLM on their phone or use the same interface with a remote model. It supports assistant conversations and character chats, with controls for how chats are structured and how models generate replies. It's open source under AGPL-3.0.
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
2.2KUpdated 3 days agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Guardrails
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
655Updated 2 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.
46Updated 5 days agoAGPL-3.0
macOS · Windows · Linux · Browser Extension#Code execution#GGUF#llama.cpp backend
Vyact is an open source desktop workspace for people who want to use a local LLM with their documents, email and code. It runs on Apple Silicon Macs, Windows and Linux x64 under the AGPL-3.0 license. Intel Macs aren't supported.