5.2KUpdated 6 days agoApache-2.0
Docker#Multimodal input
XTuner is an open-source LLM training engine for researchers and teams training large mixture-of-experts (MoE) models on their own hardware. It supports GPU and Ascend NPU training, with an emphasis on memory use and distributed training efficiency at scales reaching a trillion parameters.
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
1.7KUpdated 2 days agoApache-2.0
#Batch processing#Code execution#Multimodal input
Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.
9.2KUpdated 7 days agoApache-2.0
Docker#Image-to-image#Inpainting#Multimodal input
ModelScope combines a hosted model and dataset hub with a Python library you can run locally. It's for developers and researchers who want to use AI models in their own applications, fine-tune them on their own data, or compare their performance. The library is open source under Apache 2.0.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
1.5KUpdated 2 weeks agoApache-2.0
#Home Assistant integration#Multimodal input#Ollama integration
LLM Vision is a free, open-source Home Assistant integration for people who want their smart home to interpret what cameras see. It uses multimodal LLMs to describe images, video files, live feeds and Frigate events, then uses those results in notifications and automations.
37.5KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#LLM tracing#MCP#Multimodal input
Claude Code Router is an open-source local model gateway for developers who use coding agents and want to manage their model providers in one place. It runs on macOS, Windows and Linux, with Docker and a CLI with a browser interface also available. The project uses the MIT license.
2KUpdated 6 months agoAGPL-3.0
macOS · Windows · Linux#Code execution#LM Studio integration#MCP
Witsy is an AGPL-3.0 desktop AI assistant for macOS, Windows and Linux that connects MCP tools to local and cloud models. It's for people who want document chat, writing help and voice features in one app, with a choice of where their models run.
1KUpdated 7 days agoMIT
macOS · Windows · Linux#Batch processing#Multimodal input#Semantic search
rclip searches image folders by their visual content, so you can find photos without adding tags or importing them into a photo library. It's a local AI tool for people who keep image collections on their own computers or servers and prefer working in the terminal. It runs on Linux, Windows and Apple Silicon macOS, and it's open source under the MIT license.
4.9KUpdated 3 months ago
Linux · Docker#llama.cpp backend#Multimodal input#Ollama integration
jetson-containers is a Docker container build system for developers running local AI and robotics workloads on NVIDIA Jetson hardware. It supplies prebuilt images and lets you combine AI packages into custom containers, reducing the work of assembling compatible GPU software for JetPack/L4T.
16.2KUpdated 20 hours agoGPL-3.0
macOS · Windows · Linux#Multimodal input
LabelMe is a desktop image annotation app for people preparing computer vision datasets. It combines manual drawing with AI assistance for outlining objects and creating labels from text. It runs on 64-bit macOS, Windows and Linux..
836Updated 2 months agoMIT
Docker#LM Studio integration#Multi-user access#Multimodal input
llmcord is a self-hosted Discord bot for people who want to share LLM conversations with friends or a community. You can run the Python bot on your own machine or server, including through Docker, and connect it to local models or cloud providers. It's open source under the MIT license.
541Updated 7 months agoGPL-3.0
macOS · Windows · Linux · iOS · Android#Multimodal input#Ollama integration
Reins is an open-source chat app for people using self-hosted LLMs through Ollama. It runs on iOS, Android, macOS, Linux and Windows, giving Ollama users a mobile and desktop interface for experimenting with models. Reins is the client; Ollama provides the model backend.
14.1KUpdated 3 years ago
macOS · Windows · Linux · Web#Multimodal input#Works offline
SadTalker generates talking head videos from a single portrait and an audio recording. The project states an Apache 2.0 license and removal of its earlier noncommercial restriction. It runs locally on Windows, Linux and macOS, and suits creators who want to animate a face without recording a person on camera. Its animation includes facial expressions and head movement, with examples covering speech and singing in different languages.
14.4KUpdated 21 hours agoApache-2.0
#MCP#Multi-agent workflows#Multimodal input
LiveKit Agents is a framework for developers building voice assistants, phone agents, and apps that combine speech with video or text. Agents join LiveKit rooms as participants, so they can interact with people through web and mobile apps or telephone calls. The Apache 2.0 project lets you run the entire stack on your own servers, including the LiveKit media server.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
1.6KUpdated 9 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
677Updated 5 months agoMIT
macOS#Multimodal input#Ollama integration#OpenAI-compatible API
Obsidian Local GPT brings AI writing assistance into Obsidian, with local Ollama models for private, offline use or connections to OpenAI-compatible services. It's for people who want help with their notes while keeping control over where the AI runs. The plugin is open source under the MIT license.
49.1KUpdated 2 days agoAGPL-3.0
Linux · Docker · Web#Multi-user access#Multimodal input#OpenAI-compatible API
New API is a self-hosted AI gateway for developers and teams that want several model providers behind one service. It builds on One API and converts between OpenAI Chat Completions, Responses, Anthropic Messages and Gemini formats, so apps and agents can switch providers without changing each client's connection settings.
boltai.comChat With Your Documents
macOS#Code execution#Human approval#LM Studio integration
BoltAI is a native Mac app for people who want local LLMs and cloud AI services in the same workspace. It runs on Intel and Apple Silicon Macs and suits coding, writing, research and document analysis. The interface uses SwiftUI and AppKit.
38.7KUpdated 11 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API
Langchain-Chatchat is a self-hosted application for asking questions about your own documents and using AI agents. It focuses on Chinese-language use and open models, with a fully offline setup that can keep documents and model processing on your hardware. Its code is open source under Apache 2.0.
41.2KUpdated 1 day agoAGPL-3.0
macOS · Docker · Web#Code execution#Hybrid search#LM Studio integration
AstrBot brings AI assistants into messaging apps such as Telegram, Discord, Slack, QQ and WeCom. It's an open source platform under AGPL-3.0 for people building personal companions, customer support bots or team automation. You can run it on your own computer or server, including through Docker, or use its desktop app for browser-style chat.
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
11.1KUpdated 19 hours ago
Docker · Web#Multimodal input#Voice activity detection
TEN Framework is a self-hosted framework for developers building voice AI agents and multimodal conversational apps. It focuses on low-latency, real-time conversations and supports both RTC and WebSocket connections. You can run its agent examples locally with Docker or deploy them on your own server.
1.7KUpdated 2 days ago
Linux#Multimodal input#Quantization
RKLLM is a software stack for developers building local AI applications on Rockchip hardware. It uses the chip's neural processing unit (NPU) to run language and multimodal models on development boards, with support for the RK3588, RK3576, RK3562 and RV1126B series.
12.4KUpdated 4 weeks agoApache-2.0
macOS · Windows · Linux#Code execution#Multimodal input#Tool calling
Agent S is an open-source AI agent framework that controls ordinary desktop and web applications through the screen, mouse and keyboard. It runs on macOS, Windows and Linux and is aimed at developers, computer-use researchers and people building automation for their own desktops. You describe the task in natural language; the agent clicks, types and scrolls without requiring an API integration or a separate script for each application.
12.6KUpdated 1 week agoApache-2.0
#LM Studio integration#Multimodal input#OpenAI-compatible API
LLM is an Apache-2.0 command-line tool and Python library for sending prompts to local models and remote APIs. Local model support comes through plugins; cloud providers require their own API access. It can also connect to an arbitrary OpenAI-compatible Chat Completions endpoint, including LM Studio.
47.7KUpdated 1 month agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA
text-generation-webui, also called TextGen, runs language models on your own hardware through a desktop app or a self-hosted browser interface. It's for people who want private chat and writing tools, and developers who need a local model API. It works offline without telemetry; web search and page fetching use the internet.
6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
4.4KUpdated 2 days agoMIT
Web#LoRA#Multimodal input#Ollama integration
Ollama JavaScript connects Node.js and browser applications to models running through Ollama. It's for developers building chat interfaces, AI agents or other apps that need a local LLM backend. The library is open source under the MIT license, with TypeScript types and an API that follows Ollama's REST interface.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
1.6KUpdated 2 months agoGPL-3.0
Linux#Code execution#Multimodal input#Ollama integration
Alpaca is an open source AI chat client for people who want to run models on their own device. It uses Ollama to download and manage local models, then lets you chat without an internet connection. Conversations stay on your device in SQLite files, and local models have no direct internet access.
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
937Updated 2 years agoApache-2.0
#Hugging Face integration#Multimodal input#Works offline
Molmo is Ai2's family of vision-language models, with code for running and training models on your own hardware. It's for developers and researchers who need to work with images and text, adapt a model, or evaluate it against visual tasks. The Python codebase is open source under Apache 2.0 and builds on OLMo, adding image encoding and generative evaluation.
20.4KUpdated 3 months agoMIT
Web#Multimodal input#Tool calling
SWE-agent is an open-source AI coding agent that lets a language model use tools to attempt fixes for issues in GitHub repositories. It's aimed at developers and researchers studying how agents handle real software tasks. The project is in maintenance-only mode: mini-swe-agent has superseded it, and the maintainers recommend that successor for new users.