Tools tagged with "Streaming inference"

50 tools
Favicon of Meetily

Meetily

6 videos
A local AI meeting assistant for macOS and Windows that records and transcribes calls offline, with Ollama or your own API key for summaries.

31.3KUpdated 3 weeks agoMIT

macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference

Self-hosted AI phone and customer support agent for Windows, macOS, Linux and Docker, with SIP routing and recordings stored on your machine.

1.1KUpdated 2 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#Home Assistant integration#MCP#Streaming inference

Favicon of Pipecat

Pipecat

3 videos
Open-source Python framework for voice AI agents. Run it locally or on your servers, with WebRTC and telephony support under the BSD 2-Clause license.

16KUpdated 22 hours agoBSD-2-Clause

#LLM tracing#Multi-agent workflows#Multimodal input

Favicon of vLLM

vLLM

9 videos
An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

93KUpdated 2 hours agoApache-2.0

macOS · Docker#Batch processing#Distributed execution#GGUF

Favicon of LM Studio

LM Studio

16 videos
Download local language models, chat with documents and connect apps to a local model API on macOS, Windows or Linux.

lmstudio.aiComputer and Browser Agents

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

Favicon of llama.cpp

llama.cpp

11 videos
An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

130KUpdated 1 hour agoMIT

Web#Code execution#GGUF#Hugging Face integration

Favicon of Pydantic AI

Pydantic AI

4 videos
Open source Python AI agent SDK with typed outputs, Ollama support, and an optional self-hosted model gateway. Licensed MIT.

20.3KUpdated 2 hours agoMIT

#Human approval#LLM tracing#MCP

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

54KUpdated 2 days agoMIT

macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference

Local AI memory captures screen activity and audio on macOS, Windows, and Linux, with searchable history, on-device models, and MCP access for agents.

21.8KUpdated 21 hours ago

macOS · Windows · Linux#MCP#Multimodal input#Ollama integration

Self-hosted text-to-speech with voice cloning, multilingual speech and emotion control. Code and weights use the FISH AUDIO RESEARCH LICENSE.

32.9KUpdated 2 weeks ago

#Batch processing#Multilingual#Multimodal input

Favicon of Ollama

Ollama

31 videos
Open-source local LLM runner for macOS, Windows, Linux and Docker, with optional cloud models and coding agent integrations.

182KUpdated 17 hours agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input

Favicon of Qwen3

Qwen3

1 video
A language model family with public weights for local CPU or GPU use through Ollama, llama.cpp, and LM Studio, plus server deployment.

27.7KUpdated 9 months ago

#Batch processing#GGUF#Hugging Face integration

Favicon of VibeVoice

VibeVoice

2 videos
Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

54.5KUpdated 4 weeks agoMIT

#Hugging Face integration#Multilingual#Quantization

Offline speech-to-text app for iPhone, iPad and Apple Silicon Macs, with local speaker labels, transcript summaries on Mac and no account required.

whispernotes.appChat With Your Documents

macOS · iOS#MCP#Multilingual#Speaker diarization

A Python voice AI agent library under the MIT license, with self-hosted telephony, OpenAI and Anthropic integrations, and local speech options.

3.8KUpdated 2 years agoMIT

Docker#Streaming inference#Tool calling

Local AI voice conversion software for Windows, Linux and Apple Silicon Macs. Converts speech and singing from a short voice sample; GPL-3.0 and archived.

3.9KUpdated 1 year agoGPL-3.0

macOS · Windows · Linux · Web#Hugging Face integration#Streaming inference#Voice conversion

Self-hosted AI chat and document Q&A runs on Linux, macOS and Windows with local or cloud models. Apache 2.0 licensed; archived and no longer maintained.

12KUpdated 12 months agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access

A self-hosted text-to-speech server using Piper and Coqui XTTS v2, with voice cloning and an OpenAI-compatible API. Archived and no longer maintained.

857Updated 2 years agoAGPL-3.0

macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API

An open source LLM programming language that combines Python logic with output constraints and supports local llama.cpp and Transformers models or cloud APIs.

4.2KUpdated 1 year agoApache-2.0

Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration

Favicon of XTTS v2

XTTS v2

1 video
Local text-to-speech model with voice cloning and streaming audio, available through Coqui TTS on Linux, macOS and Windows.

2.3KUpdated 4 months agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning

An LLM chat frontend that connects to ChatGPT, Gemini, Claude, and other models through your own API keys.

typingmind.comChat and Assistants

#Agent Skills#MCP#Multimodal input

Open-source text-to-speech software for local voice cloning and streaming speech generation, with Apache 2.0 licensing and NVIDIA GPU deployment.

23.8KUpdated 4 months agoApache-2.0

Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference

Open-source text-to-speech software generates custom voices locally on NVIDIA GPUs or Apple Silicon, with Docker support and an Apache 2.0 license.

14.9KUpdated 2 years agoApache-2.0

macOS · Windows · Docker#Streaming inference#Voice cloning

A self-hosted LLM inference server under Apache 2.0, with Docker deployment, multi-GPU support and an OpenAI-compatible chat API. The project is archived.

10.9KUpdated 6 months agoApache-2.0

Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

Open-source text-to-speech built on Llama, with local inference, voice cloning and streaming audio. Uses Apache 2.0; Baseten offers cloud hosting.

6.3KUpdated 10 months agoApache-2.0

#Hugging Face integration#llama.cpp backend#LoRA

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 20 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 21 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

An open-source voice interface for text LLMs, with local hosting, Ollama and vLLM support, and speech models that require a CUDA GPU.

1.5KUpdated 3 weeks agoMIT

Docker · Web#Ollama integration#OpenAI-compatible API#Streaming inference

Favicon of Moonshine

Moonshine

3 videos
An on-device AI voice toolkit for speech recognition, intent recognition and text to speech, with support for desktop, mobile, browsers and Raspberry Pi.

11.2KUpdated 1 month ago

macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference

A self-hosted model serving library for Python and LLM APIs. Run it on a laptop or cluster with request batching, streaming, and CPU or GPU resources.

44KUpdated 18 hours agoApache-2.0

#Batch processing#Hugging Face integration#ONNX

An open-source AI music model and synthesis engine that runs locally on Apple Silicon, with a macOS app and AUv3 plugin for DAWs. Apache 2.0 licensed.

1.8KUpdated 2 months agoApache-2.0

macOS#MLX#Streaming inference

A self-hosted Discord LLM bot that connects to Ollama, LM Studio or cloud APIs, with shared conversations and an MIT open-source license.

836Updated 2 months agoMIT

Docker#LM Studio integration#Multi-user access#Multimodal input

An open-source voice AI agent framework for Python and Node.js. Run the stack on your own servers or use LiveKit Cloud. Licensed under Apache 2.0.

14.4KUpdated 20 hours agoApache-2.0

#MCP#Multi-agent workflows#Multimodal input

Local OCR model that converts images and PDFs to text or Markdown. Runs on NVIDIA GPUs with vLLM or Transformers under the MIT license.

23.9KUpdated 8 months agoMIT

Linux#Batch processing#Hugging Face integration#Multimodal input

An open-source voice AI framework that processes speech directly, with local inference on Mac and iPhone through MLX and self-hosted server backends.

11.2KUpdated 5 months agoApache-2.0

macOS · iOS · Web#Hugging Face integration#MLX#Quantization

A self-hosted text-to-speech API for Kokoro-82M. Generate speech locally on CPU, NVIDIA GPU or Apple Silicon, with multi-speaker audio and captions.

5.5KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API