Tools tagged with "OpenAI-compatible API"

100+ tools
An AI agent CLI that runs locally, connects to Ollama or hosted providers, and can read, search and edit files in your workspace. Node.js 24+ required.

3.5KUpdated 3 weeks ago

#Ollama integration#OpenAI-compatible API#Tool calling

A self-hosted AI chatbot platform that connects messaging apps to Ollama, LM Studio or cloud models, with agent tools and document knowledge bases.

41.2KUpdated 1 day agoAGPL-3.0

macOS · Docker · Web#Code execution#Hybrid search#LM Studio integration

A self-hosted LLM API server that runs ExLlamaV3 models on your hardware, with OpenAI-compatible endpoints and an AGPL-3.0 license.

1.4KUpdated 2 days agoAGPL-3.0

Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

Open-source AI data processing framework for local machines and Ray clusters, with multimodal cleaning, deduplication and Apache 2.0 licensing.

7.1KUpdated 2 days agoApache-2.0

Docker#Batch processing#Distributed execution#Multimodal input

An open-source AI agent plugin for self-hosted Mattermost that connects to Ollama, vLLM or cloud providers to search and summarize team conversations.

251Updated 1 day agoApache-2.0

#Human approval#MCP#Multi-user access

Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 1 day agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API

A CLI and Python library for local LLMs and remote APIs, with SQLite prompt history, embeddings, and support for OpenAI-compatible endpoints.

12.6KUpdated 1 week agoApache-2.0

#LM Studio integration#Multimodal input#OpenAI-compatible API

Local LLM desktop app for Windows, macOS and Linux. Run GGUF models offline or connect other apps through OpenAI- and Anthropic-compatible APIs.

47.7KUpdated 1 month agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#GGUF#llama.cpp backend#LoRA

A self-hosted text-to-speech API for Kokoro-82M. Generate speech locally on CPU, NVIDIA GPU or Apple Silicon, with multi-speaker audio and captions.

5.5KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API

An open-source AI email assistant for Thunderbird that connects to local models through Ollama or LM Studio, or to cloud services such as ChatGPT and Claude.

347Updated 2 days agoGPL-3.0

#Batch processing#LM Studio integration#Multilingual

An open-source local LLM training framework with LoRA, multimodal support, and deployment through vLLM, SGLang or LMDeploy. Apache 2.0 licensed.

15.8KUpdated 2 days agoApache-2.0

Web#Distributed execution#Hugging Face integration#LoRA

Favicon of Continue

Continue

4 videos
An open-source coding agent available as a CLI, VS Code extension, and JetBrains plugin, with an Apache 2.0 license and no active maintenance.

36.1KUpdated 2 months agoApache-2.0

VS Code · JetBrains#Code execution#Human approval#MCP

Favicon of Alpaca

Alpaca

1 video
A local AI chat client for Ollama models that works offline after download and can also connect to cloud models through OpenAI-compatible APIs.

1.6KUpdated 2 months agoGPL-3.0

Linux#Code execution#Multimodal input#Ollama integration

Python library for running GGUF models locally through llama.cpp, with a self-hosted OpenAI-compatible server and CPU or GPU support.

10.6KUpdated 1 week agoMIT

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

Favicon of koboldcpp

koboldcpp

1 video
Local LLM runner for GGUF and GGML models on Windows, macOS and Linux, with CPU or GPU support, a browser UI and an AGPL-3.0 license.

11.9KUpdated 4 days agoAGPL-3.0

macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

An open-source AI browser automation tool with a no-code builder, a Playwright-compatible SDK, and support for Ollama and cloud models.

23.1KUpdated 1 hour agoAGPL-3.0

#Code execution#MCP#Multi-agent workflows

An open-source LLM CLI for macOS, Linux and Windows that connects to Ollama or cloud providers and supports document chat, shell commands and AI agents.

10.5KUpdated 7 months agoApache-2.0

macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input

Favicon of Aider

Aider

1 video
An open-source terminal coding assistant that edits code with local or cloud LLMs, maps repositories, and records changes in Git. Licensed under Apache 2.0.

49.3KUpdated 4 months agoApache-2.0

macOS · Windows · Linux#Code execution#Git integration#Multimodal input

A self-hosted RSS reader that ranks news using AI tags and your scoring rules. Runs with Docker and supports OpenAI, Gemini and compatible model APIs.

395Updated 4 weeks ago

Docker · Web#Multi-user access#OpenAI-compatible API

An open source RLHF framework for training models on your own NVIDIA GPUs, with HuggingFace model support and Ray, vLLM and DeepSpeed backends.

10.1KUpdated 2 weeks agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

An open-source local LLM compiler and deployment engine with GPU support across desktop, browser and mobile platforms, plus an OpenAI-compatible API.

23.2KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API

A local LLM inference engine with OpenAI and Anthropic-compatible APIs. Runs on macOS, Linux and Windows with CPU, CUDA or Apple Silicon support.

7.7KUpdated 5 days agoMIT

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration

An open-source Obsidian AI plugin for searching and writing notes with Ollama, LM Studio, Claude Code or Codex. Local indexes stay on your device.

7.8KUpdated 2 days agoAGPL-3.0

#LM Studio integration#Multi-agent workflows#Ollama integration

Favicon of Elia

Elia

1 video
A terminal chat client for local LLMs through Ollama or LocalAI and cloud models from OpenAI and Anthropic, with locally stored conversations.

2.5KUpdated 2 years agoApache-2.0

#Ollama integration#OpenAI-compatible API

An open-source local AI API under Apache 2.0 that connects to Ollama, llama.cpp and other OpenAI-compatible servers for document retrieval and agent workflows.

57.6KUpdated 1 week agoApache-2.0

Docker · Web#Code execution#llama.cpp backend#MCP

Favicon of Lemonade

Lemonade

1 video
An open source local AI server for chat, image generation, and speech on Windows, macOS, and Linux, with APIs for apps and agents.

5.8KUpdated 2 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend

Self-hosted AI research assistant for web, papers and private documents. Runs on Windows, macOS and Linux with Ollama or cloud models. MIT licensed.

9.1KUpdated 23 hours agoMIT

macOS · Windows · Linux · Docker · Web#llama.cpp backend#MCP#Multi-user access

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 3 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.7KUpdated 22 hours ago

Docker#Batch processing#Distributed execution#LoRA

Self-hosted speech API for transcription, translation and speech generation. Runs via Docker on CPU or GPU with faster-whisper, Kokoro and Piper.

3.7KUpdated 5 months agoMIT

Docker#OpenAI-compatible API#Streaming inference

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

Favicon of llama-swap

llama-swap

1 video
A local AI proxy that switches models on demand through OpenAI and Anthropic compatible APIs. Runs on macOS, Windows, Linux and FreeBSD under MIT.

5.8KUpdated 2 days agoMIT

macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend

An open-source Android LLM chat app that runs GGUF models on-device through llama.cpp or connects to Ollama, OpenAI and Claude. Licensed under AGPL-3.0.

2.8KUpdated 1 week agoAGPL-3.0

Android#GGUF#llama.cpp backend#Ollama integration

Favicon of GPT4All

GPT4All

1 video
An open-source local AI chatbot for Windows, macOS and Linux. Run models without a GPU or cloud API, and chat privately with your documents.

77.4KUpdated 1 year agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API

Self-hosted subtitle generator runs Whisper locally on CPU or NVIDIA GPU and connects to Bazarr, Plex, Jellyfin, Emby and Tautulli. MIT licensed.

1.5KUpdated 2 months agoMIT

Docker#Batch processing#Multilingual#OpenAI-compatible API

Local document converter turns PDFs and Office files into Markdown, JSON or HTML, with OCR on CPU, NVIDIA GPUs or Apple Silicon and optional LLM support.

40.1KUpdated 3 weeks agoApache-2.0

macOS · Linux · Web#Batch processing#llama.cpp backend#Multilingual