15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
3.4KUpdated 23 hours agoApache-2.0
Docker · Web#Multilingual#Multimodal input
MTEB is an Apache 2.0 Python toolkit for evaluating embedding models and retrieval systems. It runs evaluations through Python or a command-line interface and publishes an interactive leaderboard.
7.7KUpdated 2 years agoMIT
Docker · Web#Multilingual
MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.
goodsnooze.gumroad.comDictation and Voice Typing
macOS · iOS#Batch processing#Multilingual#Ollama integration
MacWhisper is a native macOS transcription app for people working with interviews, lectures, meetings and other recorded audio. It runs speech recognition on your own Mac, so local transcription keeps audio on your device. It also offers cloud transcription through services such as OpenAI, ElevenLabs and Deepgram, which send audio off your machine.
23.8KUpdated 20 hours agoMPL-2.0
macOS · Windows · Linux · iOS · Android#Multilingual#Persistent memory#RAG
Brave Leo is an AI assistant built into the Brave browser, with Bring Your Own Model support for people who want to use their own local or remote models while browsing. It can work with third-party APIs as well as Brave's hosted model choices. The browser runs on macOS, Windows, Linux, Android, and iOS.
30KUpdated 10 months agoApache-2.0
Windows#Multilingual
EasyOCR is a Python OCR library for developers who want to extract text from images on their own hardware. It reads text in photographs and dense documents, so it can serve both scene-text recognition and document processing. It's open source under Apache 2.0.
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
1.1KUpdated 2 years agoGPL-3.0
macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization
WhisperWriter turns microphone speech into text and types it into the window you're working in. It's for people who want voice input in their existing desktop apps, with a choice between transcription on their own computer and an external service. The Python app runs on Windows, macOS and Linux and uses the GPL-3.0 open-source license.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
214Updated 3 weeks agoMIT
Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual
Wyoming Piper connects Piper's local text-to-speech engine to Home Assistant and other clients that use the Wyoming protocol. It's for people building a voice assistant on their own hardware who need speech generation as a self-hosted service. The project is open source under the MIT license.
1.8KUpdated 20 hours agoApache-2.0
Linux · Docker#Distributed execution#Hugging Face integration#Multilingual
NeMo Curator is an open source Python toolkit for ML engineers and data teams preparing AI training datasets on their own hardware. It handles text, images, video and audio, with reusable pipelines that can run on a laptop or scale across a multi-node Ray cluster. NVIDIA uses it to prepare data for Nemotron models.
immersivetranslate.comOCR and Document Scanning
macOS · iOS · Android · Browser Extension#Inpainting#Multilingual
Immersive Translate is an AI translation extension and mobile app that keeps original text alongside its translation. It's aimed at students, researchers and people who read foreign-language material for work. The bilingual page layout lets readers compare passages without replacing the source text.
enconvo.comAI Workflow Automation
macOS · iOS#LM Studio integration#MCP#MLX
Enconvo is a native AI assistant and agent for Mac that can use your screen and selected text as context, then work inside your apps. It's for people who want help with writing, research and everyday tasks alongside the app they're using. A sidebar keeps the agent beside the current app, while text selection tools give quick access to editing and translation.
1.4KUpdated 3 days ago
#GGUF#Home Assistant integration#Hugging Face integration
Home-LLM connects Home Assistant to language models running on your own hardware, so you can control smart devices through voice or chat. It's for Home Assistant users who want natural language control without relying on a cloud service or subscription. The project pairs a custom integration with small models trained specifically for smart home commands.
181Updated 7 months ago
macOS · Windows · Linux#Multilingual#Ollama integration#OpenAI-compatible API
LocalWriter brings local LLM writing assistance into LibreOffice Writer for people who want to draft and revise text inside their documents. It runs on macOS, Windows and Linux and connects to a separate model runner, including Ollama and text-generation-webui. With a backend on your own machine, text processing stays local.
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
37.7KUpdated 1 year agoMIT
#Multilingual#Voice cloning
OpenVoice is an open-source voice cloning tool that uses a short recording to reproduce a speaker's voice in generated speech. It's for developers and creators who need a recognizable voice across languages, with control over how that voice sounds. The Python project is MIT licensed for commercial use.
1.6KUpdated 9 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.
10KUpdated 2 years agoAGPL-3.0
#Hugging Face integration#Multilingual
PDF-Extract-Kit is a local AI model toolbox for developers and researchers building document processing applications. It extracts text, tables and mathematical formulas from PDFs, with separate models for identifying page elements and recognizing their contents. It's open source under AGPL-3.0, written in Python, and supports CPU or GPU execution on your own hardware.
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.
3KUpdated 3 weeks agoMIT
macOS#Batch processing#Hugging Face integration#Multilingual
OpenSuperWhisper is a local speech-to-text app for people who want to dictate or transcribe recordings on an Apple Silicon Mac. It supports Whisper and Parakeet, with model downloads available inside the app. The project is open source under the MIT license.
2.9KUpdated 9 months agoApache-2.0
Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization
Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.
5.5KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API
Kokoro-FastAPI runs the Kokoro-82M speech model on your own machine or server and exposes an OpenAI-compatible speech API. It's for developers adding local text-to-speech to assistants, reading apps or audiobook workflows. Speech generation runs locally, and the API doesn't require an OpenAI account.
6.6KUpdated 1 year ago
Windows · Linux · Web#Batch processing#Inpainting#Multilingual
MuseTalk is a local AI lip-sync model for creators and developers working on video dubbing or virtual avatars. It edits the face in an existing video to match supplied speech, including Chinese, English and Japanese audio. It runs on Windows and Linux with NVIDIA GPUs, and can process videos generated by MuseV.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
347Updated 2 days agoGPL-3.0
#Batch processing#LM Studio integration#Multilingual
ThunderAI brings AI writing and email processing into Thunderbird for people who want help with their inbox while choosing where their messages go. It can use local models through Ollama or an OpenAI-compatible server such as LM Studio. Cloud connections send the selected content to ChatGPT, the OpenAI API, Google Gemini or Claude instead.
9.1KUpdated 1 year agoApache-2.0
macOS · Windows#Batch processing#Multilingual#ONNX
Kokoro is a text-to-speech model and inference library for developers who want to generate speech on their own hardware or servers. Its compact Kokoro-82M model suits personal projects and production applications, with Apache 2.0 licensing for both the library and model weights.
10.8KUpdated 8 months agoMIT
Windows · Docker · Web#Multi-user access#Multilingual
Doccano is a self-hosted text annotation tool for machine learning practitioners who need labeled training or evaluation data. It runs on your own machine or server, with a browser interface and Docker support. The software is open source under the MIT license.
76.8KUpdated 3 days agoApache-2.0
Windows#Multilingual
Tesseract is an open source OCR engine for extracting text from images, with a command line program and a library developers can embed in their own applications. It's suited to document processing workflows and software that needs text recognition. The project uses the Apache 2.0 license and doesn't include a graphical app.
superwhisper.comDictation and Voice Typing
macOS · Windows · iOS · Android#Multilingual#Works offline
Superwhisper is an AI dictation app for macOS, Windows, iOS and Android that turns speech into text in the app you're using. It's for people who prefer speaking to typing, including developers dictating requests to coding assistants. Speech recognition can run locally and offline, or use cloud models for remote processing.
4.7KUpdated 5 days agoMIT
#Batch processing#Multilingual#Quantization
CTranslate2 is an open-source C++ and Python library for developers running Transformer models on their own hardware or servers. It handles translation, text generation, text encoding and speech recognition. Its custom runtime focuses on reducing inference time and memory use compared with general-purpose deep learning frameworks.
6.5KUpdated 2 months agoMIT
Linux#Multilingual#Works offline
Argos Translate runs neural machine translation locally, so you can translate text without sending it to an online service. It's for people who need offline translation and developers who want to add it to their own software. The Python library and desktop interface use the same translation engine, and the project is open source under the MIT license.
24.3KUpdated 4 days agoBSD-2-Clause
macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual
WhisperX is an open source speech-to-text tool for people transcribing interviews, meetings, and long recordings on their own computer. It builds on OpenAI's Whisper to produce transcripts with word-level timestamps and optional speaker labels.