25.8KUpdated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Hybrid search#llama.cpp backend#Multi-user access
kotaemon is a self-hosted document chat app for people who want to ask questions across their files and check where the answers came from. It runs in a browser on Windows, macOS or Linux, with Docker also supported. The project uses the Apache 2.0 license.
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
23.1KUpdated 1 hour agoAGPL-3.0
#Code execution#MCP#Multi-agent workflows
Skyvern uses vision models to carry out tasks in a browser, such as filling forms, collecting data, and working through logins. It's for teams automating websites whose layouts change, as well as developers who want AI actions alongside Playwright. Instead of depending on fixed page selectors, it identifies visible elements and decides how to interact with them.
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
10.5KUpdated 7 months agoApache-2.0
macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input
aichat brings Ollama and cloud AI services into the same terminal interface for developers and people who work at the command line. It runs locally on macOS, Linux and Windows, with Android support through Termux. Model processing happens through the backend you choose: Ollama supports local models, while providers such as OpenAI, Claude and Gemini process requests in the cloud.
49.3KUpdated 4 months agoApache-2.0
macOS · Windows · Linux#Code execution#Git integration#Multimodal input
Aider is an Apache 2.0 open-source coding assistant for developers who work in a terminal and keep their projects in Git. It can help start a project or make changes to an existing codebase. You can use local LLMs or connect to cloud models from providers including Anthropic, DeepSeek and OpenAI.
17.3KUpdated 11 months agoApache-2.0
Windows · Linux · Web#Hugging Face integration#Multimodal input
FramePack is an open source desktop app for making videos from a still image and a written motion prompt. It runs on Windows and Linux, with generation handled by your own NVIDIA GPU. It suits people who want to make AI video locally and see the clip develop as it renders.
28.6KUpdated 1 day agoMIT
macOS · Windows · Linux#LM Studio integration#MCP#Multi-agent workflows
Semantic Kernel is an MIT-licensed SDK for developers adding AI agents to their applications. It supports local models through Ollama, LMStudio and ONNX, alongside cloud services such as OpenAI and Azure OpenAI. You choose the model backend. The SDK runs on Windows, macOS and Linux and supports C#, Python and Java.
23.7KUpdated 1 day agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
verl is a Python library for teams training large language models on their own GPU infrastructure. It's the open-source implementation of HybridFlow, aimed at researchers and engineers who need reinforcement learning after initial model training. It uses the Apache 2.0 license.
26.6KUpdated 6 days agoGPL-3.0
macOS · Linux · Docker#Hybrid search#Multimodal input#RAG
Typesense combines typo-tolerant site search with vector and semantic search in a self-hosted engine. It's for developers building searchable apps, product catalogs or AI search over their own data. The C++ engine uses an in-memory architecture for low-latency results as users type.
10.1KUpdated 2 weeks agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.
2.9KUpdated 24 hours agoMIT
Web · VS Code#Code execution#Hugging Face integration#MCP
Inspect AI is a Python framework for researchers and developers testing language models and AI agents. Developed by the UK AI Security Institute and Meridian Labs, it evaluates coding, reasoning, knowledge, behavior and multimodal understanding, including tasks where agents must take actions to succeed.
6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
41.9KUpdated 6 days agoGPL-3.0
macOS · Windows · Linux · iOS · Android · Web#MCP#Multimodal input#Ollama integration
Chatbox is an AI chat client for people who want local models and cloud providers in the same app. It connects to Ollama for local LLM use and supports GPT, Claude, Gemini, Grok and DeepSeek with your own API keys. Chatbox also offers its own hosted model service.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
91Updated 1 day agoAGPL-3.0
Web#Multi-user access#Multilingual#Multimodal input
Nextcloud Assistant brings AI into Nextcloud Hub's documents, email, chat and calendar. It's for teams that want help with shared work while choosing where AI processing happens. With an on-premises model, data stays on your server. The app is open source under AGPL-3.0.
13.7KUpdated 11 months agoMIT
Linux · Web#Hugging Face integration#Multimodal input
TRELLIS is a local AI model for generating 3D assets from images or text prompts, aimed at 3D artists and researchers exploring asset creation. It can produce meshes, radiance fields and 3D Gaussians from the same underlying representation, so you can choose an output suited to your rendering or editing work.
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
2.5KUpdated 1 day ago
macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input
Roboflow Inference is a self-hosted computer vision server for teams building camera and image analysis systems. It runs on your own computer, server, or edge device and combines model predictions with workflows for tracking, counting, measuring, and responding to events. Roboflow also offers hosted servers and a Serverless Cloud API, where processing runs on its infrastructure.
5.8KUpdated 2 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
13KUpdated 23 hours agoApache-2.0
Docker#Agent Skills#Hugging Face integration#Knowledge graphs
txtai is a Python framework for developers building search applications, chat with their data, and AI agents on their own hardware or servers. Its embeddings database combines sparse and dense vector search with graphs and relational data, so the same system can find related content and supply context to language models. It's open source under Apache 2.0.
10.6KUpdated 2 days agoApache-2.0
Linux · Web#Hugging Face integration#Multilingual#Multimodal input
YuE is an open-source music generation project for musicians, songwriters and developers who want to turn lyrics and a style prompt into songs with vocals and accompaniment. Its YuE2 models create an editable melody and chord score before generating the recording, so you can review the composition and change musical details before hearing the result.
38.4KUpdated 4 days agoMIT
#Code execution#MCP#Multimodal input
DSPy is a Python framework for developers building AI applications whose tasks need clear inputs, predictable output types, and measurable results. You define what a language model should produce, then compose those tasks into a larger program. It's open source under the MIT license.
7.7KUpdated 12 months ago
#Hugging Face integration#Multimodal input#Quantization
Llama is Meta's family of large language models for developers, researchers and businesses that want to run models on their own hardware or servers. Its downloadable weights let you build generative AI applications with local inference. Access requires license acceptance and approval, and the weights use custom licensing for research and commercial use.
18.5KUpdated 1 day agoApache-2.0
#LLM tracing#Multimodal input
DeepEval is a Python framework for testing AI agents, RAG pipelines, and chatbots in your own environment. It's for developers and ML teams who need to compare models or prompts and catch quality regressions before deployment. The open-source framework uses the Apache 2.0 license and fits into Pytest, Python scripts, notebooks, and CI/CD.
28.4KUpdated 1 day agoApache-2.0
macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input
Label Studio is a self-hosted platform for teams preparing training data or evaluating AI outputs through human review. It handles text, images, audio, video and time series in the same application, including tasks that combine several data types. The open source edition uses the Apache 2.0 license and runs locally or on your own server, with Docker deployment and browser access. A separate hosted cloud edition runs on the provider's infrastructure.
8.1KUpdated 3 days agoApache-2.0
#Batch processing#Distributed execution#Hugging Face integration
LMDeploy is an open-source toolkit for developers serving language and vision-language models on their own hardware. It combines model compression with inference and self-hosted APIs, so teams can use it for batch processing or as the model backend for an application. It uses the Apache 2.0 license.
39.2KUpdated 7 days agoApache-2.0
macOS · Windows · Web#Multimodal input
UI-TARS Desktop is an open-source AI agent that controls computer interfaces through natural language requests. It runs on Windows and macOS and supports browser use, with operators for both local and remote computers. It's for people who want an agent to carry out tasks in existing apps, such as changing a VS Code setting or checking an issue on GitHub.
14.7KUpdated 22 hours ago
Docker#Batch processing#Distributed execution#LoRA
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
12.1KUpdated 4 months agoApache-2.0
Docker#Multimodal input#Structured output
Jina Reader turns web pages and documents into text that LLMs can use, with Markdown or JSON output. It's for developers building AI agents, search tools and systems that answer questions using retrieved documents. You can self-host the Apache 2.0 service code in Docker or use Jina's hosted API.
10.2KUpdated 1 year agoMIT
#Hugging Face integration#Multimodal input
InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
19.5KUpdated 1 week agoApache-2.0
Docker#LoRA#Multimodal input#Prompt caching
KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.
59.4KUpdated 1 day ago
#Hybrid search#MCP#Multilingual
Meilisearch combines keyword search and AI retrieval in a search engine you can host on your own server. It's for developers building search into websites, applications, product catalogs, or internal data tools. Meilisearch Cloud provides a separate, fully managed hosted service.