21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
10.6KUpdated 2 days agoMIT
#Batch processing#Ollama integration#Streaming inference
Ollama Python connects Python applications to models running through Ollama on your own machine or to Ollama's cloud service. It's for developers adding local LLM features to scripts, chat applications, or other Python projects. The library is open source under the MIT license and requires a running Ollama service for local use.
2.9KUpdated 6 months agoMIT
macOS · Docker#Batch processing#Hugging Face integration#Multimodal input
Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
47.6KUpdated 10 months agoMIT
Windows · Linux#Batch processing#Multilingual#Works offline
Umi-OCR turns screenshots, images, and scanned documents into text on Windows and Linux. It works offline. The software is free under the MIT License, and recognition runs locally with built-in language libraries. It suits people who need to copy text from a screen, process a folder of images, or make scans searchable without sending them to an online OCR service.
50KUpdated 6 months agoAGPL-3.0
macOS · Windows · Linux#Batch processing
Upscayl enlarges low-resolution photos and graphics with AI. Its free, open source desktop app runs locally on Linux, macOS, and Windows, making it an option for people who want to improve images on their own computer. It needs a Vulkan-compatible GPU and is licensed under AGPL-3.0.
25.6KUpdated 3 hours agoMIT
#Batch processing#Hugging Face integration#Quantization
faster-whisper is a Python library for people building local speech transcription into their own software. It runs OpenAI's Whisper models through CTranslate2, with faster processing and lower memory use than the original Whisper implementation in the project's comparisons. It runs on a CPU. NVIDIA GPUs are supported too, and the code is open source under the MIT license.
89Updated 23 hours agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Hugging Face integration#ONNX
PixlStash is an open-source image manager for photographers, AI creators and people curating image datasets. It combines library search with tools for reviewing tags, ranking images and sending selected work through ComfyUI. You can use a desktop app or a self-hosted server with a browser interface.