Tools tagged with "Multimodal input"

200+ tools
Favicon of smolagents

smolagents

1 video
An open-source Python AI agent library with local models through Transformers or Ollama, cloud API support, and Docker sandboxing.

29.6KUpdated 1 week agoApache-2.0

#Code execution#Hugging Face integration#MCP

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 2 days agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

An open-source visual editor for React websites. Run it locally or self-host it, with AI chat and support for Next.js and Tailwind CSS.

26.8KUpdated 2 months agoApache-2.0

Web#Code execution#Git integration#Multimodal input

An open-source Ruby AI framework for Ruby and Rails apps, with Ollama, hosted providers and OpenAI-compatible endpoints under an MIT license.

4.4KUpdated 1 day agoMIT

#Human approval#Multi-agent workflows#Multimodal input

A self-hosted AI workspace with parallel model chats and answer merging. Connect Ollama, LM Studio or cloud providers using your own API keys.

7.1KUpdated 1 day agoMIT

Docker · Web#LM Studio integration#Multimodal input#Ollama integration

.NET library for local LLM apps using Ollama, with streaming chat, embeddings and model management. Open source under MIT; also supports Ollama cloud.

1.4KUpdated 2 months agoMIT

#Multimodal input#Ollama integration#Streaming inference

Self-hosted AI agent platform with document Q&A, workflows and MCP tools. Connect local DeepSeek, Llama or Qwen models, or use cloud providers.

22.9KUpdated 3 days agoGPL-3.0

Docker · Web#MCP#Multimodal input#RAG

A self-hosted AI coding assistant for full-stack web apps, with Ollama and LM Studio support, desktop apps, and MIT-licensed source code.

19.9KUpdated 8 months agoMIT

macOS · Windows · Linux · Docker · Web#Code execution#Git integration#LM Studio integration

An open source Neovim AI coding assistant with Ollama support, cloud model connections, inline edits, and agents such as Claude Code and Codex.

6.9KUpdated 22 hours agoApache-2.0

#Agent Client Protocol#Human approval#MCP

Favicon of Axolotl

Axolotl

1 video
Apache 2.0 LLM fine-tuning framework for local or cloud GPUs, supporting NVIDIA, AMD, LoRA, QLoRA and multimodal training.

12.5KUpdated 2 days agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

An open-source AI chat app for macOS, iOS and visionOS that connects to your own Ollama server and stores conversation history on your device.

6KUpdated 3 months agoApache-2.0

macOS · iOS#Multimodal input#Ollama integration#Works offline

An open-source LLM client that brings local models through Ollama and llama.cpp, plus cloud services such as Claude and Gemini, into Emacs.

3.5KUpdated 6 days agoGPL-3.0

#Git integration#Human approval#llama.cpp backend

A self-hosted inference framework that coordinates NVIDIA GPU clusters with vLLM, SGLang or TensorRT-LLM and exposes an OpenAI-compatible API.

8.2KUpdated 21 hours ago

#Distributed execution#Multimodal input#OpenAI-compatible API

A self-hosted OCR model that extracts text and page structure from PDFs and images, with vLLM, Hugging Face Transformers and CPU inference support.

9.2KUpdated 6 months agoMIT

Docker#Hugging Face integration#Multilingual#Multimodal input

A Python research assistant that answers questions with citations from local documents. Supports self-hosted models through LiteLLM; licensed under Apache 2.0.

9.3KUpdated 2 months agoApache-2.0

#Multilingual#Multimodal input#RAG

Favicon of OpenVINO

OpenVINO

2 videos
Open-source AI inference toolkit for local or self-hosted deployment on Linux, Windows and macOS, with CPU, Intel GPU and NPU support.

10.9KUpdated 1 day agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX

A local AI workbench for macOS, Windows and Linux. Evaluate agents, optimize prompts and run fully offline with Ollama or use cloud APIs.

5.1KUpdated 22 hours ago

macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows

Open-source AI podcast generator in Python with local HuggingFace models for transcripts and cloud speech services for multilingual audio.

6.6KUpdated 5 months agoApache-2.0

Docker · Web#Hugging Face integration#Multilingual#Multimodal input

AI video generation model you can run on your own GPUs, with text-to-video and image-to-video support, ComfyUI integration and downloadable weights.

12.6KUpdated 3 months ago

Web#Multimodal input#Quantization

Open-source video-to-audio and text-to-audio software you can run locally on a GPU. Uses an MIT license; tested on Ubuntu.

2.3KUpdated 7 months agoMIT

Linux · Web#Multimodal input

An AI agent workspace with hosted and Docker self-hosted options for building agents, scheduling tasks, and sharing projects with a team.

82.9KUpdated 4 hours ago

Linux · Docker · Web#MCP#Multi-agent workflows#Multi-user access

A self-hosted AI video generation model with ComfyUI and Diffusers support, image animation, video editing, and fine-tuning tools.

11KUpdated 9 months agoApache-2.0

#Hugging Face integration#LoRA#Multimodal input

Python library for running and training embedding and reranker models locally, with Apache 2.0 licensing and pretrained models on Hugging Face.

19.1KUpdated 1 week agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

Open-source Rust ML framework for running models locally on CPUs, NVIDIA GPUs or in browsers, with Apache 2.0 licensing and quantized LLM support.

21.1KUpdated 2 days agoApache-2.0

macOS · Web#GGUF#Hugging Face integration#Multilingual

Self-hosted embedding and reranking API with MIT licensing, Hugging Face models, and CPU, NVIDIA, AMD and Apple MPS support.

2.9KUpdated 6 months agoMIT

macOS · Docker#Batch processing#Hugging Face integration#Multimodal input

A self-hosted AI answer engine that cites web sources and works with Ollama, local model servers, or cloud models.

36.9KUpdated 4 weeks agoMIT

Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API

Local AI app for Android and iOS. Run Gemma 4 on-device, ask questions about photos, transcribe audio and compare model performance.

24.8KUpdated 22 hours agoApache-2.0

iOS · Android#Agent Skills#Hugging Face integration#Multilingual

Search your documents, organize notes and schedule tasks with a self-hosted AI assistant that connects to local text and vision models.

922Updated 5 months agoMIT

macOS · Windows · Linux · Docker · Web#llama.cpp backend#LM Studio integration#MLX