31.3KUpdated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference
Meetily is a local AI meeting assistant for people who want meeting notes while keeping recordings on their own device. It captures calls from Zoom, Google Meet, Microsoft Teams and other meeting software without placing a bot in the meeting. You can watch the transcript appear during the call, then generate a summary.
1.1KUpdated 2 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Home Assistant integration#MCP#Streaming inference
Tel-Agent is a self-hosted AI receptionist for businesses that want phone calls and customer messages handled in one place. It runs on your own hardware and keeps recordings and searchable transcripts on your machine. It's free and open source under the AGPL-3.0 license.
16KUpdated 22 hours agoBSD-2-Clause
#LLM tracing#Multi-agent workflows#Multimodal input
Pipecat is a Python framework for developers building conversational AI agents that handle speech, video, text, and images. You can run it on your own machine or servers, wherever Python runs. It's open source under the BSD 2-Clause license.
93KUpdated 2 hours agoApache-2.0
macOS · Docker#Batch processing#Distributed execution#GGUF
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
lmstudio.aiComputer and Browser Agents
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
LM Studio is a desktop application for downloading and running language models on macOS, Windows and Linux. You can search for models, manage downloads and chat with them through the app. Downloaded models can run offline, including document chat that uses files on your computer.
130KUpdated 1 hour agoMIT
Web#Code execution#GGUF#Hugging Face integration
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
20.3KUpdated 2 hours agoMIT
#Human approval#LLM tracing#MCP
Pydantic AI is a Python SDK for developers building AI agents into their own applications. Its main draw is Pydantic validation across agent tools and results, so an agent can return structured data that application code can check and use. The SDK is MIT licensed.
54KUpdated 2 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
21.8KUpdated 21 hours ago
macOS · Windows · Linux#MCP#Multimodal input#Ollama integration
Screenpipe records screen activity and audio as searchable local history that AI agents can use as context. It's for people who want to recall past work and teams that want agents to draft follow-ups or update work records using what actually happened. Raw history stays on your device by default.
32.9KUpdated 2 weeks ago
#Batch processing#Multilingual#Multimodal input
Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.
182KUpdated 17 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
27.7KUpdated 9 months ago
#Batch processing#GGUF#Hugging Face integration
Qwen3 is a family of language models from Alibaba Cloud’s Qwen team for people who want to run models locally or on their own servers. It spans smaller and larger dense models as well as mixture-of-experts models. The weights are publicly available.
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
whispernotes.appChat With Your Documents
macOS · iOS#MCP#Multilingual#Speaker diarization
Whisper Notes is an offline transcription app for iPhone, iPad and Apple Silicon Macs. It's for people recording interviews, lectures or meetings who need the audio and transcripts to stay on their device. It requires no account and has no cloud sync, analytics or tracking.
3.8KUpdated 2 years agoMIT
Docker#Streaming inference#Tool calling
Vocode is an open source Python library for developers building voice AI agents, with a self-hosted telephony server and support for live conversations through a computer's microphone and speakers. It connects speech recognition, an LLM, and speech synthesis in one library. The code uses the MIT license.
3.9KUpdated 1 year agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Streaming inference#Voice conversion
Seed-VC changes recorded speech or singing to sound like a voice supplied in a short reference clip, without training a separate model for that speaker. It runs locally on Windows, Linux and Apple Silicon Macs, with uses in audio production, live streaming and online meetings. The project is archived and no longer maintained.
12KUpdated 12 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access
h2oGPT is a self-hosted ChatGPT alternative for people who want to chat with local models and ask questions about their own documents. The project is archived and no longer maintained. It's open source under Apache 2.0, with support for Linux, macOS, Windows and Docker.
857Updated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API
OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
4.2KUpdated 1 year agoApache-2.0
Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration
LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
typingmind.comChat and Assistants
#Agent Skills#MCP#Multimodal input
TypingMind is a browser chat frontend for people who want to choose their model providers and manage conversations in one place. It connects to cloud APIs such as OpenAI, Claude and Gemini, and supports custom endpoints for locally hosted models such as Ollama and LocalAI. The frontend does not run model inference itself.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
14.9KUpdated 2 years agoApache-2.0
macOS · Windows · Docker#Streaming inference#Voice cloning
Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.
10.9KUpdated 6 months agoApache-2.0
Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Text Generation Inference (TGI) is a self-hosted LLM server for developers and teams serving models through an API on their own hardware. The repository is archived; its README describes maintenance mode and recommends other inference engines for new deployments. Its focus is handling concurrent generation requests and making efficient use of GPU memory.
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
3.5KUpdated 20 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
5.1KUpdated 21 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
1.5KUpdated 3 weeks agoMIT
Docker · Web#Ollama integration#OpenAI-compatible API#Streaming inference
Unmute adds spoken conversation to text LLMs using Kyutai's speech recognition and speech synthesis models. It's for developers who want a self-hosted voice interface while keeping their choice of language model. The project uses the MIT license, and a hosted browser demo is available at Unmute.sh.
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
44KUpdated 18 hours agoApache-2.0
#Batch processing#Hugging Face integration#ONNX
Ray Serve is a self-hosted Python library for developers building inference APIs that combine models with application logic. It runs on a laptop, on-premise servers, Kubernetes, or cloud infrastructure you choose. It's open source under Apache 2.0.
1.8KUpdated 2 months agoApache-2.0
macOS#MLX#Streaming inference
Magenta RealTime 2 is a local AI music model and synthesis engine for musicians and developers who want to play or build AI musical instruments on a laptop. It generates streaming audio in real time, with open weights and code under the Apache 2.0 license.
836Updated 2 months agoMIT
Docker#LM Studio integration#Multi-user access#Multimodal input
llmcord is a self-hosted Discord bot for people who want to share LLM conversations with friends or a community. You can run the Python bot on your own machine or server, including through Docker, and connect it to local models or cloud providers. It's open source under the MIT license.
14.4KUpdated 20 hours agoApache-2.0
#MCP#Multi-agent workflows#Multimodal input
LiveKit Agents is a framework for developers building voice assistants, phone agents, and apps that combine speech with video or text. Agents join LiveKit rooms as participants, so they can interact with people through web and mobile apps or telephone calls. The Apache 2.0 project lets you run the entire stack on your own servers, including the LiveKit media server.
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
11.2KUpdated 5 months agoApache-2.0
macOS · iOS · Web#Hugging Face integration#MLX#Quantization
Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.
5.5KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API
Kokoro-FastAPI runs the Kokoro-82M speech model on your own machine or server and exposes an OpenAI-compatible speech API. It's for developers adding local text-to-speech to assistants, reading apps or audiobook workflows. Speech generation runs locally, and the API doesn't require an OpenAI account.