Tools tagged with "Multimodal input"

200+ tools
A self-hosted multimodal retrieval engine for AI apps, with ColPali visual search, Docker deployment, and MCP access through Claude or Open Web UI.

3.7KUpdated 5 days ago

Docker · Web#MCP#Multi-user access#Multimodal input

A self-hosted vector search engine for text and images, with built-in embedding generation, Docker deployment, and an Apache 2.0 license.

5KUpdated 6 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#RAG

A self-hostable AI chat interface with Docker deployment, local model connections through LocalAI and RWKV-Runner, and desktop clients.

88.8KUpdated 2 months agoMIT

macOS · Windows · Linux · Docker · Web#MCP#Multimodal input#OpenAI-compatible API

A local LLM app that runs models offline on iOS and macOS, supports text and vision models, and uses ggml and llama.cpp under the MIT license.

2.1KUpdated 8 months agoMIT

macOS · iOS#llama.cpp backend#Multimodal input#RAG

A self-hosted AI retrieval system with document search, cited answers and research agents. Runs in Python or Docker and uses the MIT license.

8KUpdated 11 months agoMIT

Docker#Hybrid search#Knowledge graphs#Multi-user access

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

Vision-language models that answer image questions, read text in images and detect objects, with downloadable weights and GPU research code.

3.5KUpdated 1 year agoApache-2.0

#Hugging Face integration#Multimodal input

A self-hosted vision-language model project for image chat, with a local Gradio interface, GPU inference and Apache 2.0 code.

25KUpdated 2 years agoApache-2.0

macOS · Web#LoRA#Multimodal input#Quantization

An open-source AI model family for on-premises deployment, licensed under Apache 2.0, with language, speech, vision and guardrail models.

273Updated 2 years agoApache-2.0

Linux#Guardrails#Hugging Face integration#LM Studio integration

An open-source image and video generation framework with 4K text-to-image models, laptop GPU support, and ComfyUI and Diffusers integrations.

9.2KUpdated 2 weeks agoApache-2.0

#ControlNet#LoRA#Multimodal input

An open-source image and text model for local image classification without task-specific training. Runs through PyTorch on CPU or CUDA GPUs under MIT.

34.4KUpdated 6 months agoMIT

#Batch processing#Multimodal input

A local image captioning model with open weights, Apache 2.0 code, SFW and NSFW coverage, and support for ComfyUI and vLLM.

1.3KUpdated 7 months agoApache-2.0

Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API

An open-source object detection model that finds objects using text prompts. Run it locally with PyTorch on a CPU or NVIDIA GPU. Apache 2.0 licensed.

10.6KUpdated 2 years agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input

A Python library for running and training CLIP image-text models on your own hardware, with local checkpoints and Hugging Face model support.

14.2KUpdated 5 days ago

#Hugging Face integration#Multimodal input

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

8.9KUpdated 8 months agoApache-2.0

Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration

An open-source image generation model that runs locally, edits images, and uses multiple references. Supports Diffusers and carries an MIT license.

4.3KUpdated 10 months agoMIT

Web#Hugging Face integration#Image-to-image#LoRA

A local multimodal AI model for answering image questions and generating pictures, with downloadable weights and a Gradio interface.

17.8KUpdated 2 years agoMIT

Web#Batch processing#Hugging Face integration#Multimodal input

An open-source ChatGPT alternative for Chrome that uses Ollama models locally to chat across tabs, analyze files, and assist with writing.

1.1KUpdated 7 months agoAGPL-3.0

Web · Browser Extension#Multilingual#Multimodal input#Ollama integration

Open-source text-to-speech model for local English dialogue generation, with voice cloning, NVIDIA GPU inference and an Apache 2.0 license.

19.4KUpdated 10 months agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input#Voice cloning

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 20 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

Local AI voice assistant for Linux and Windows with camera vision, persistent memory and MCP tools. Uses Ollama or cloud APIs. MIT licensed.

5.7KUpdated 2 weeks agoMIT

macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows

A browser-based LLM chat app that connects to Ollama and OpenAI servers, stores chat data in your browser, and supports text, vision and reasoning models.

1.2KUpdated 12 months agoMIT

macOS · Windows · Linux · Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API

An open-source terminal AI agent that runs on your machine with Ollama or cloud models, edits files, executes code, and supports persistent tasks.

4.4KUpdated 20 hours agoMIT

macOS · Windows · Linux · Web · JetBrains#Agent Client Protocol#Code execution#Git integration

Local AI lip-sync software matches video to new speech, with pretrained models and training code for research, academic and personal use.

13.2KUpdated 1 year ago

Linux · Docker#Multilingual#Multimodal input

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 22 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

Local AI portrait animation software turns images and audio into talking-head videos, with editable facial landmarks and Apache 2.0 source code.

4.3KUpdated 6 months agoApache-2.0

Linux · Web#Hugging Face integration#Multimodal input

Self-hosted AI chat interface under Apache 2.0. Connect it to Ollama, llama.cpp or cloud APIs, with optional model routing and MCP tools.

11KUpdated 1 day agoApache-2.0

Docker · Web#llama.cpp backend#MCP#Multi-user access

A desktop image dataset editor with local AI captioning on CPU or NVIDIA GPU. Runs on Windows, Linux and macOS under GPL-3.0.

1.4KUpdated 12 months agoGPL-3.0

macOS · Windows · Linux#Batch processing#Multimodal input

A Python toolkit for local LLM compression and inference on Linux, macOS and Windows, with GPTQ, AWQ, GGUF and integrations for vLLM and SGLang.

1.3KUpdated 1 day ago

macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA

A text-to-music model with melody conditioning and local GPU inference. AudioCraft code is MIT licensed; pretrained weights have a noncommercial license.

23.7KUpdated 2 years agoMIT

#Multimodal input

An open-source Python toolkit for evaluating embedding models and retrieval systems, with Sentence Transformers support and an interactive leaderboard.

3.4KUpdated 23 hours agoApache-2.0

Docker · Web#Multilingual#Multimodal input

A local AI image generator for macOS that runs Stable Diffusion and FLUX.2 Klein offline on Apple Silicon. Open source under GPL-3.0.

8KUpdated 5 days agoGPL-3.0

macOS#ControlNet#Image-to-image#Multimodal input

A self-hosted AI add-on for paperless-ngx that extracts text and organizes documents using local Ollama models or cloud APIs. Open source under MIT.

2.7KUpdated 6 days agoMIT

Docker · Web#Batch processing#Human approval#Multimodal input

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 21 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

Open-source OCR library for Node.js and Python that converts documents to Markdown using cloud vision models. MIT licensed; page images leave your machine.

12.3KUpdated 1 year agoMIT

Linux#Multimodal input#Structured output

A desktop AI chat app that stores data locally, connects to your own model providers, and supports files, web search and MCP tools.

chatwise.appChat With Your Documents

#MCP#Multimodal input#Tool calling