Tools tagged with "Multimodal input"

200+ tools
A coding assistant CLI that writes and runs code from plain-language requests. Runs locally or in Docker, with local models or OpenAI and Anthropic APIs.

55.1KUpdated 2 years agoMIT

Windows · Docker#Code execution#Multimodal input

AI content moderation models check text and images in LLM inputs and outputs. Use downloadable models or Meta's hosted Llama API.

4.4KUpdated 21 hours ago

#Guardrails#Multilingual#Multimodal input

AI writing assistant for macOS that works with Apple Foundation Models, Ollama and LM Studio, or cloud providers through your own API keys.

kerlig.comChat With Your Documents

macOS#LM Studio integration#MCP#Multimodal input

A paid Windows AI agent app that coordinates custom agents, stores data locally, and supports Ollama alongside ChatGPT and Mistral AI.

nurgo-software.comAI Workflow Automation

Windows#Code execution#Multi-agent workflows#Multimodal input

An open-source AI music generator that runs locally on macOS, Windows and Linux, with text or audio style prompts and Apache 2.0 code and DiT weights.

2.3KUpdated 10 months agoApache-2.0

macOS · Windows · Linux · Docker#Hugging Face integration#Multimodal input

A local document retrieval library that matches text queries to page images without OCR. MIT-licensed Python code supports NVIDIA and Apple Silicon GPUs.

2.8KUpdated 1 month agoMIT

macOS#Batch processing#Hugging Face integration#LoRA

Jina’s embedding models encode multilingual text and media for retrieval, with local weights, noncommercial licenses and commercial deployment options.

jina.aiEmbedding and Reranker Models

Docker#GGUF#LoRA#MLX

An AI companion that talks, edits files and plays out stories. Run models on your own hardware or use cloud services, with Windows and Linux support.

voxta.aiAI Characters and Roleplay

Windows · Linux · Android · Web#Code execution#MCP#Multimodal input

Open-source AI training framework built on PyTorch. Train on local CPUs or GPUs, fine-tune HuggingFace models, and serve models on your own server.

11.8KUpdated 4 days agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

Local LLM quantization library for smaller model weights and inference on your hardware. MIT licensed, with CPU and GPU support; archived and unmaintained.

2.3KUpdated 1 year agoMIT

Linux#Batch processing#GGUF#Hugging Face integration

Higgs Audio V2, now Higgs TTS 2, is a downloadable speech model for expressive narration, multilingual dialogue and voice cloning.

8.4KUpdated 4 months agoApache-2.0

#Batch processing#Hugging Face integration#Multilingual

A local AI companion for Windows, macOS and Linux with voice chat, Live2D avatars and camera input. Runs offline with local models or connects to cloud APIs.

14KUpdated 5 months ago

macOS · Windows · Linux · Web#GGUF#LM Studio integration#MCP

An open-source speech generation model that uses text and audio context, runs on a CUDA-compatible GPU, and integrates with Hugging Face Transformers.

14.7KUpdated 1 year agoApache-2.0

Windows#Hugging Face integration#Multimodal input

An open source AI character interface you can run locally, with voice chat, VRM avatars, and support for Ollama, llama.cpp, and cloud APIs.

1.6KUpdated 1 year agoMIT

Windows · Docker · Web#llama.cpp backend#LM Studio integration#Multimodal input

Open-source AI video generator that runs locally with miniFLUX or SD3 models, supports Apple Silicon, and includes a browser interface.

3.2KUpdated 2 years agoMIT

macOS · Web#Hugging Face integration#Multimodal input

An open-source audio AI model for local transcription, translation and Q&A. Runs offline with vLLM or Transformers under Apache 2.0.

10.8KUpdated 3 months agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

An open-source AI animation module for Stable Diffusion that runs locally, works with personalized models, and supports image or sketch guidance.

12.3KUpdated 2 years agoApache-2.0

Web#Hugging Face integration#LoRA#Multimodal input

An open-source AI roleplay chat app you can self-host with Docker, with custom characters, group chats and connections to Kobold, Claude and OpenRouter.

784Updated 4 months agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#Multi-user access#Multimodal input#Persistent memory

Self-hosted LLM gateway with evaluation, A/B testing, and Ollama support. Open source under Apache 2.0; archived and no longer maintained.

11.7KUpdated 4 months agoApache-2.0

Docker · Web#Batch processing#LLM tracing#Multimodal input

An LLM chat frontend that connects to ChatGPT, Gemini, Claude, and other models through your own API keys.

typingmind.comChat and Assistants

#Agent Skills#MCP#Multimodal input

Self-hosted RAG framework for document Q&A, with Docker, Ollama and Infinity support. Apache 2.0 licensed; archived and no longer maintained.

4.4KUpdated 7 months agoApache-2.0

Docker · Web#Batch processing#Multimodal input#Ollama integration

Open source AI video generator you can run on your own GPUs, with image and text inputs, model training tools, and an Apache 2.0 license.

29.9KUpdated 6 months agoApache-2.0

#Hugging Face integration#Multimodal input

Favicon of SkyReels

SkyReels

1 video
AI video generator with a hosted web app for synchronized sound and reference control, plus downloadable models for local GPU inference.

7.6KUpdated 8 months ago

Web#Hugging Face integration#Multimodal input

Local AI video generator with downloadable models, NVIDIA GPU support and LoRA fine-tuning. Code and the 2B model use Apache 2.0.

13KUpdated 11 months agoApache-2.0

Windows · Web#Hugging Face integration#LoRA#Multimodal input

Self-hosted AI application server with an OpenAI-compatible API, local Ollama and vLLM backends, document search and agent tool calling. MIT licensed.

8.4KUpdated 21 hours agoMIT

#Agent Skills#Batch processing#Guardrails

Local LLM web interface for Windows, macOS and Linux. Use GGUF models, Ollama or cloud APIs, with local chat storage and an Apache 2.0 license.

4.8KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

Self-hosted LLM inference engine for Hugging Face models, with OpenAI-compatible APIs, multimodal support, and CPU or GPU execution under AGPL-3.0.

1.9KUpdated 3 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

Favicon of Qwen-Image

Qwen-Image

1 video
An open-source image generation and editing model for local deployment, with Chinese text rendering, multi-image edits and an Apache 2.0 license.

8.4KUpdated 8 months agoApache-2.0

Web#Image-to-image#LoRA#Multimodal input

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

A local vision-language model for image analysis, document extraction and video understanding, with Apache 2.0 weights and Hugging Face Transformers support.

huggingface.coOCR and Document Scanning

Linux#Batch processing#Hugging Face integration#Multimodal input

Multimodal AI models you can run locally with Mistral's GPU inference library, which is open source under Apache 2.0 and archived.

10.8KUpdated 3 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#Tool calling

Open-source vision language model for local image and text tasks, with Apache 2.0 licensing, Transformers support, and a small GPU memory footprint.

3.9KUpdated 1 week agoApache-2.0

#Hugging Face integration#LoRA#Multimodal input

An open source on-device AI engine for local LLMs and image models, with iOS, Android, CPU and GPU support under Apache 2.0.

16.2KUpdated 1 day agoApache-2.0

Windows · iOS · Android#Image-to-image#Multimodal input#ONNX

A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

2.1KUpdated 3 weeks agoApache-2.0

Linux#GGUF#Guardrails#Hugging Face integration

LLM models for local or cloud inference, with downloadable weights and function calling. The Apache 2.0 inference library is archived.

10.8KUpdated 3 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#Tool calling

An open-source portrait animation tool that runs on Ubuntu with an NVIDIA GPU, turning a still image and English speech into a talking video.

8.7KUpdated 2 years agoMIT

Linux#Hugging Face integration#Multimodal input#ONNX