18.5KUpdated 1 day agoApache-2.0
#LLM tracing#Multimodal input
DeepEval is a Python framework for testing AI agents, RAG pipelines, and chatbots in your own environment. It's for developers and ML teams who need to compare models or prompts and catch quality regressions before deployment. The open-source framework uses the Apache 2.0 license and fits into Pytest, Python scripts, notebooks, and CI/CD.
36.2KUpdated 7 days agoMIT
#Knowledge graphs#RAG#Semantic search
GraphRAG builds a knowledge graph from text so an LLM can answer questions that depend on connections across documents or themes across a whole collection. It's for developers and researchers working with private datasets, such as business documents, proprietary research, or communications.
4.6KUpdated 1 week agoApache-2.0
Linux
NVIDIA Container Toolkit lets Docker containers use NVIDIA GPUs on a Linux machine or server. It's for developers and server operators who need GPU acceleration for containerized workloads, including self-hosted AI software. The project is open source under the Apache 2.0 license.
28.4KUpdated 1 day agoApache-2.0
macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input
Label Studio is a self-hosted platform for teams preparing training data or evaluating AI outputs through human review. It handles text, images, audio, video and time series in the same application, including tasks that combine several data types. The open source edition uses the Apache 2.0 license and runs locally or on your own server, with Docker deployment and browser access. A separate hosted cloud edition runs on the provider's infrastructure.
8.7KUpdated 7 months agoMIT
macOS · Windows · Linux#Multilingual
Audiblez turns EPUB e-books into M4B audiobooks using Kokoro-82M text-to-speech on your own computer. It's for readers who want spoken versions of their books and control over the narrator, reading speed and sections included. The app is open source under the MIT license.
8.1KUpdated 3 days agoApache-2.0
#Batch processing#Distributed execution#Hugging Face integration
LMDeploy is an open-source toolkit for developers serving language and vision-language models on their own hardware. It combines model compression with inference and self-hosted APIs, so teams can use it for batch processing or as the model backend for an application. It uses the Apache 2.0 license.
322Updated 2 weeks agoMIT
iOS · Android · Docker · Web#Batch processing#Code execution#Guardrails
Tiledesk is a self-hosted platform for building AI agents and connecting them to human support teams. It's aimed at businesses automating customer conversations, internal information searches and workflows through a visual builder. You can run it on your own server with Docker or Kubernetes, or use its hosted cloud service.
27.9KUpdated 1 day agoApache-2.0
#MCP#Structured output#Tool calling
FastMCP is an open-source Python framework for developers connecting AI agents to their own tools and data through the Model Context Protocol (MCP). It supports locally running servers and connections to remote servers, with server development and client access in the same framework. It uses the Apache 2.0 license.
39.2KUpdated 7 days agoApache-2.0
macOS · Windows · Web#Multimodal input
UI-TARS Desktop is an open-source AI agent that controls computer interfaces through natural language requests. It runs on Windows and macOS and supports browser use, with operators for both local and remote computers. It's for people who want an agent to carry out tasks in existing apps, such as changing a VS Code setting or checking an issue on GitHub.
14.7KUpdated 22 hours ago
Docker#Batch processing#Distributed execution#LoRA
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
91.9KUpdated 1 year agoMIT
#Hugging Face integration
DeepSeek-V3 and DeepSeek-R1 are downloadable language models for developers who want to run text generation and reasoning on their own hardware. V3 is a mixture-of-experts text model. R1 builds on DeepSeek-V3-Base and focuses on reasoning tasks such as math and coding. The V3 and R1 repositories describe their respective models; the R1 repository links to downloadable weights.
12.1KUpdated 4 months agoApache-2.0
Docker#Multimodal input#Structured output
Jina Reader turns web pages and documents into text that LLMs can use, with Markdown or JSON output. It's for developers building AI agents, search tools and systems that answer questions using retrieved documents. You can self-host the Apache 2.0 service code in Docker or use Jina's hosted API.
3.7KUpdated 5 months agoMIT
Docker#OpenAI-compatible API#Streaming inference
Speaches is a self-hosted speech server for developers who want transcription, translation and speech generation on their own hardware. Its OpenAI-compatible API lets applications use local speech models through tools and SDKs built for OpenAI's API. The project is open source under the MIT license.
13.8KUpdated 23 hours agoApache-2.0
#Semantic search
OpenSearch combines document and enterprise search with vector retrieval for AI applications. It's for developers building search into their products and teams analyzing application logs, infrastructure performance, or security events. The suite uses the Apache 2.0 license throughout, including its data ingestion and dashboard components.
37.7KUpdated 2 days agoApache-2.0
macOS · Windows · Linux · Docker#Code execution#LM Studio integration#MCP
Playwright MCP lets AI agents control a browser by reading structured page information rather than interpreting screenshots. It's for developers building browser automation, exploratory tests or agent workflows that need to keep a browser session open across repeated actions. The server runs locally on macOS, Windows and Linux, or as a self-hosted service.
798Updated 1 month agoApache-2.0
Linux
NVIDIA DCGM monitors and manages NVIDIA data-center GPUs on your own Linux servers. It's for infrastructure teams running GPU clusters, including those hosting AI workloads, who need to track hardware health, investigate slow jobs and control power use. It supports x86_64 and aarch64 (SBSA) systems.
3.2KUpdated 1 month agoAGPL-3.0
macOS · Windows · Linux#Inpainting#LoRA
OneTrainer is an open-source application for training diffusion models on your own machine, with dataset preparation and model previews in the same interface. It's for people adapting image or video models with their own training data. It runs on Windows, macOS and Linux under the AGPL-3.0 license.
10.2KUpdated 1 year agoMIT
#Hugging Face integration#Multimodal input
InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.
8.5KUpdated 4 weeks agoMIT
macOS · Windows · Linux#LoRA#Quantization
bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
8.2KUpdated 2 weeks agoGPL-3.0
macOS#Image-to-image#Inpainting#Visual workflows
Dream Textures is an open-source Blender add-on for artists who want to generate images and apply AI textures to 3D work inside Blender. It runs Stable Diffusion on your own machine and connects image generation to scene depth, texture projection, and animation rendering.
5.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend
llama-swap is a self-hosted proxy for people running several AI models on their own hardware. It starts the model server a request needs and swaps out another when necessary, so you don't have to keep every model loaded or manage separate API connections in your apps.
6KUpdated 2 days ago
#Home Assistant integration
Scrypted is a server-based video integration platform for people who want to view their home cameras through their existing smart home apps and displays. It supports most cameras and focuses on low latency streaming. Its NVR plugin adds continuous recording and smart detections for people who also need a record of what happened.
19.5KUpdated 1 week agoApache-2.0
Docker#LoRA#Multimodal input#Prompt caching
KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.
33.5KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#GGUF#ONNX
Netron displays neural network and machine learning models as visual graphs. Developers and researchers can inspect a model's structure through desktop apps for macOS, Windows and Linux, or through the browser viewer at netron.app.
23.8KUpdated 11 months ago
Docker#Code execution#RAG
PandasAI is a Python library for people who want to ask questions about their data in plain language. It works with SQL databases, CSV and parquet data, and can answer questions across multiple pandas DataFrames. Developers can use it in their own applications; analysts can use conversational queries to reduce the code they write for individual questions.
18KUpdated 1 day ago
Docker#Distributed execution#Hugging Face integration#Quantization
Megatron-LM is a Python framework for research teams training large language models on NVIDIA GPU infrastructure. It pairs ready-made training scripts with Megatron Core, a library developers can use to build their own training systems. Its focus is distributed training, with benchmarks on H100 clusters spanning thousands of GPUs.
2.8KUpdated 1 week agoAGPL-3.0
Android#GGUF#llama.cpp backend#Ollama integration
ChatterUI is an Android chat app for people who want to run a local LLM on their phone or use the same interface with a remote model. It supports assistant conversations and character chats, with controls for how chats are structured and how models generate replies. It's open source under AGPL-3.0.
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
1.5KUpdated 2 months agoMIT
Docker#Batch processing#Multilingual#OpenAI-compatible API
subgen generates subtitles on your own hardware for personal media libraries, including films and shows that don't have usable subtitles available. It's an open source, MIT-licensed Python service that runs in Docker or as a standalone application. Speech recognition runs locally using Whisper models through faster-whisper and stable-ts, with support for CPU processing and NVIDIA GPUs through CUDA.
59.4KUpdated 1 day ago
#Hybrid search#MCP#Multilingual
Meilisearch combines keyword search and AI retrieval in a search engine you can host on your own server. It's for developers building search into websites, applications, product catalogs, or internal data tools. Meilisearch Cloud provides a separate, fully managed hosted service.
40.1KUpdated 3 weeks agoApache-2.0
macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual
Marker is a local document converter for developers and teams turning PDFs, scans and Office files into structured text. It preserves tables, equations and page structure for document processing and AI workflows. Its pipeline reads embedded PDF text and uses Surya OCR where text is missing or damaged, rather than reading every page through a vision model.
29.6KUpdated 1 week agoApache-2.0
#Code execution#Hugging Face integration#MCP
smolagents is an open-source Python library for developers building AI agents that carry out tasks by writing and executing Python. Its CodeAgent can combine tool calls with loops, conditionals, and calculations in one action. This approach suits tasks that need several operations, rather than a single model response.
8.2KUpdated 4 months agoApache-2.0
macOS · Windows · Linux · Web#Semantic search
sqlite-vec adds vector storage and similarity search to SQLite, so developers can keep embeddings alongside application data in a local database. It's for applications that need to find related items by vector distance without running a separate vector database server. The extension is small, written in C and has no dependencies.
59.2KUpdated 1 day agoMIT
#MCP#Multi-agent workflows#Ollama integration
CrewAI is an open-source Python framework for developers building workflows with multiple AI agents on their own hardware or servers. It supports local models through Ollama and defaults to the OpenAI API for model requests. The framework uses the MIT license; a separate commercial platform provides managed deployment and governance, with on-premise and cloud options.
52.4KUpdated 2 days agoMIT
#Ollama integration#RAG#Reranking
LlamaIndex is an MIT-licensed Python framework for developers building AI agents and apps that answer questions using their own data. It connects documents and other sources to language models, then helps an app find the relevant material when a user asks something. Its open source framework can work with models served through Ollama.
2.5KUpdated 2 days agoMIT
macOS · Linux#Hugging Face integration#Multilingual
LightEval is a Python toolkit from Hugging Face for evaluating LLMs running on your own hardware or through remote services. It's for developers and researchers comparing models, investigating failures, or testing performance on tasks relevant to their work. It can evaluate a model already loaded in memory as well as one served through an endpoint.
51.1KUpdated 1 day agoMIT
Supervision is an MIT-licensed Python library from Roboflow for developers building computer vision applications around their own models. It handles the work around predictions: drawing results on images and video, following objects across frames, and turning detections into counts. It can work with images and datasets on your machine.
29.8KUpdated 1 day ago
Docker · Web#LLM tracing#MCP#Multi-user access
FastGPT is a self-hosted AI agent builder for teams that want assistants to answer questions using company documents and carry out business workflows. Its visual editor connects model calls, knowledge retrieval and tools into applications for customer support, internal knowledge search and document review. You can run the platform on your own server through Docker or use the vendor's hosted service.
15.9KUpdated 7 months agoApache-2.0
Ragas is an open-source Python library for developers who need repeatable evaluations of LLM applications and retrieval-augmented generation (RAG) systems. It combines model-based scoring with traditional metrics so teams can compare application changes using test results rather than manual judgments alone. Its license is Apache 2.0.
3.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.
16.9KUpdated 1 day ago
Docker#Hybrid search#RAG#Reranking
Weaviate is a self-hosted vector database for developers building search applications, RAG systems, recommendation engines, and chatbots. It stores data objects alongside their vector embeddings, so applications can search by meaning and filter results using structured data. You can run the database locally with Docker, deploy it on Kubernetes, or use the hosted Weaviate Cloud service.
12.5KUpdated 1 month agoApache-2.0
Web#LLM tracing#Multi-user access#Streaming inference
Chainlit is an open-source Python framework for developers building conversational AI apps with their own application logic. You can run the app on your own server and give users a browser chat interface. It's a framework for creating an app, rather than a ready-made chatbot with a fixed model.
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.
2.6KUpdated 1 week agoAGPL-3.0
Linux · Docker
jetson-stats monitors NVIDIA Jetson hardware and gives you control over its power and cooling settings. It runs locally on the board, with a terminal interface called jtop for people developing or running workloads on Jetson devices, including local AI applications.
7.7KUpdated 4 weeks agoMIT
macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration
Vibe is an open source desktop app for people who need transcripts or subtitles without uploading their recordings to a transcription service. It runs on macOS, Windows and Linux under the MIT license. Audio transcription works fully offline, with processing on your own computer.
31.2KUpdated 22 hours agoApache-2.0
Docker · Web#Knowledge graphs#MCP#Multi-user access
Cognee gives AI agents persistent memory across sessions, connecting documents, code, and conversations in a searchable knowledge graph. It's for developers who want agents to retain project context and teams whose knowledge sits across tickets, discussions, and repositories. The Python package is open source under Apache 2.0.
1.3KUpdated 2 days agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.