1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
huggingface.coOCR and Document Scanning
Linux#Batch processing#Hugging Face integration#Multimodal input
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
14.1KUpdated 2 weeks agoMIT
macOS#Batch processing#GGUF#Hugging Face integration
lm-evaluation-harness lets researchers and model developers compare language models using shared academic benchmarks and public prompts. It runs evaluations against local models and benchmarks, or sends requests to a hosted model API. EleutherAI's Python framework is open source under the MIT license and powers Hugging Face's Open LLM Leaderboard.
4.3KUpdated 4 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android#Batch processing#Semantic search
USearch is an open-source similarity search library for developers building semantic search, recommendation systems, and other applications that compare vectors. It runs within your application on your own hardware or server. Its compact C++ core supports custom definitions of similarity, including comparisons between combined image and text embeddings or geospatial data.
34.4KUpdated 6 months agoMIT
#Batch processing#Multimodal input
CLIP is an open-source image and text model that lets developers and researchers classify images using labels written in natural language, without collecting training examples for each task. It runs locally through PyTorch on a CPU or CUDA GPU. The code and model weights use the MIT license.
23.3KUpdated 1 year agoApache-2.0
macOS · Windows · Web#Batch processing#Image-to-image#Inpainting
IOPaint is a free, self-hosted AI image editor for people who want to remove objects, replace parts of a picture or extend it beyond its original edges on their own hardware. The project is archived and no longer maintained. It's open source under Apache 2.0, with a browser interface and support for CPU, GPU and Apple Silicon hardware.
4.6KUpdated 7 months agoMIT
Windows · Linux#Batch processing#Quantization#Speculative decoding
ExLlamaV2 is a local LLM inference library for developers and people hosting models on their own consumer GPUs. ExLlamaV2 is archived and no longer maintained; development continues in ExLlamaV3. The V2 library is free and open source under the MIT license, runs on Windows and Linux, and uses NVIDIA GPUs through CUDA. It supports multiple GPUs.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
17.8KUpdated 2 years agoMIT
Web#Batch processing#Hugging Face integration#Multimodal input
Janus-Pro is a multimodal AI model from DeepSeek that answers questions about images and creates pictures from text prompts. It runs on your own hardware and suits developers and researchers who want both capabilities in one model. A local Gradio demo provides a browser interface, while Hugging Face hosts a separate online demo.
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
2.8KUpdated 9 months agoApache-2.0
Windows · Linux#Batch processing#ONNX#Voice activity detection
openWakeWord is a Python library for developers building voice interfaces that listen locally for a chosen word or phrase. It includes English models for triggers such as "hey jarvis" and "alexa", plus phrases for weather and timers. The code uses Apache 2.0. Included pretrained models use CC-BY-NC-SA-4.0, which restricts commercial use.
752Updated 3 weeks agoApache-2.0
Linux · Docker · Web · Browser Extension#Batch processing#OpenAI-compatible API#Quantization
Meme Search is a free, self-hosted web app for people who want to find memes in their own collection by image content or text. It runs locally through Docker and uses AI descriptions to make images searchable, even when their filenames aren't useful. The code uses the Apache 2.0 license.
1.4KUpdated 12 months agoGPL-3.0
macOS · Windows · Linux#Batch processing#Multimodal input
TagGUI is a desktop app for people preparing image datasets for generative AI training. It combines local AI captioning with manual tag editing, so you can generate descriptions and correct them in the same workspace. It's open source under GPL-3.0 and runs on Windows, Linux and macOS, though macOS doesn't have a packaged release.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
2.7KUpdated 6 days agoMIT
Docker · Web#Batch processing#Human approval#Multimodal input
Paperless-GPT adds AI text extraction and document organization to an existing paperless-ngx library. It runs in Docker on your own server and suits people who want less manual sorting of scanned paperwork. The project is open source under the MIT license.
goodsnooze.gumroad.comDictation and Voice Typing
macOS · iOS#Batch processing#Multilingual#Ollama integration
MacWhisper is a native macOS transcription app for people working with interviews, lectures, meetings and other recorded audio. It runs speech recognition on your own Mac, so local transcription keeps audio on your device. It also offers cloud transcription through services such as OpenAI, ElevenLabs and Deepgram, which send audio off your machine.
5.1KUpdated 21 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
1.7KUpdated 2 days agoApache-2.0
#Batch processing#Code execution#Multimodal input
Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
44KUpdated 19 hours agoApache-2.0
#Batch processing#Hugging Face integration#ONNX
Ray Serve is a self-hosted Python library for developers building inference APIs that combine models with application logic. It runs on a laptop, on-premise servers, Kubernetes, or cloud infrastructure you choose. It's open source under Apache 2.0.
34.4KUpdated 2 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#MCP
ChatDev is a self-hosted platform for building teams of AI agents through a visual workflow editor. It's for people who want agents to collaborate on research, data analysis or software projects without writing the orchestration code themselves. You define each agent's role and how information passes between them.
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
1KUpdated 7 days agoMIT
macOS · Windows · Linux#Batch processing#Multimodal input#Semantic search
rclip searches image folders by their visual content, so you can find photos without adding tags or importing them into a photo library. It's a local AI tool for people who keep image collections on their own computers or servers and prefer working in the terminal. It runs on Linux, Windows and Apple Silicon macOS, and it's open source under the MIT license.
9.8KUpdated 5 months agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Hugging Face integration
PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.
8.2KUpdated 2 years ago
#Batch processing#GGUF#Hugging Face integration
GOT-OCR2.0 is an OCR model for developers and researchers who want to extract text from images on their own hardware. It handles both plain text and formatted output through a single model, with recognition modes for selected regions and documents spanning multiple pages. The Python codebase builds on Vary.
3.6KUpdated 1 day agoMIT
#Batch processing#LLM tracing#MCP
TruLens is an open-source Python tool for developers who need to find why an AI agent gives a wrong answer or spends too much on a task. It pairs step-level traces with evaluation scores, so you can connect failures to retrieval, reasoning or tool calls. It uses the MIT license and can write results to a database you run.
18.5KUpdated 1 day agoApache-2.0
Linux · Docker#Batch processing#Hugging Face integration
Parakeet is NVIDIA's speech recognition model family. The linked parakeet-tdt-0.6b-v2 is its English speech-to-text model for developers and researchers building transcription services, subtitles or voice applications. It runs locally through NeMo on Linux, with NVIDIA GPUs recommended for inference. It's a model you can embed in an application, rather than a desktop transcription app.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
4.1KUpdated 2 years agoMIT
#Batch processing#Hugging Face integration
Distil-Whisper is a family of local speech recognition models for developers building English transcription into their apps or services. It reduces Whisper's size and processing time while retaining much of its transcription accuracy. It supports English only.
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
3KUpdated 3 weeks agoMIT
macOS#Batch processing#Hugging Face integration#Multilingual
OpenSuperWhisper is a local speech-to-text app for people who want to dictate or transcribe recordings on an Apple Silicon Mac. It supports Whisper and Parakeet, with model downloads available inside the app. The project is open source under the MIT license.
1.4KUpdated 2 days agoAGPL-3.0
Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
TabbyAPI is a self-hosted LLM API server built around ExLlamaV3, for people who want local model inference behind an OpenAI-compatible API. It's the official server for that backend. The project targets personal use and small groups, and its maintainers explicitly advise against using it for production workloads.
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
4.7KUpdated 1 day agoApache-2.0
#Batch processing#Distributed execution#OpenAI-compatible API
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
10.5KUpdated 3 weeks ago
macOS · Windows · Linux · Web#Batch processing#ControlNet#Image-to-image
Easy Diffusion runs Stable Diffusion on your own computer through a browser interface. It's for people who want to generate and edit images locally without assembling the software components themselves. The free distribution bundles the required software and works on Windows, Linux and macOS.