Tools tagged with "llama.cpp backend"

73 tools
Favicon of Graphiti

Graphiti

1 video
A self-hosted Python framework for AI agent memory that tracks changing facts. Apache 2.0 licensed, with support for local LLMs and cloud APIs.

31.3KUpdated 1 day agoApache-2.0

Docker#Hybrid search#Knowledge graphs#llama.cpp backend

A terminal AI coding assistant that runs on macOS, Linux, Windows and BSD, with support for Ollama, llama.cpp and cloud model APIs.

28.4KUpdated 1 day ago

macOS · Windows · Linux · Android#llama.cpp backend#LM Studio integration#MCP

Favicon of LM Studio

LM Studio

16 videos
Download local language models, chat with documents and connect apps to a local model API on macOS, Windows or Linux.

lmstudio.aiComputer and Browser Agents

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

Favicon of MinerU

MinerU

1 video
Local document parser converts PDFs, Office files and images into Markdown, JSON and LaTeX on macOS, Windows and Linux, with optional remote parsing.

80.8KUpdated 2 days ago

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

Favicon of Ollama

Ollama

31 videos
Open-source local LLM runner for macOS, Windows, Linux and Docker, with optional cloud models and coding agent integrations.

182KUpdated 17 hours agoMIT

macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input

Favicon of Qwen3

Qwen3

1 video
A language model family with public weights for local CPU or GPU use through Ollama, llama.cpp, and LM Studio, plus server deployment.

27.7KUpdated 9 months ago

#Batch processing#GGUF#Hugging Face integration

A self-hosted AI runtime with an OpenAI-compatible API. It runs models on CPUs or GPUs and keeps inference on your own hardware.

49.3KUpdated 2 hours agoMIT

macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend

Favicon of Zed

Zed

2 videos
A native code editor for macOS, Linux, and Windows that combines parallel AI agents with shared editing, chat, and screen sharing.

91.1KUpdated 21 hours ago

macOS · Windows · Linux#Agent Client Protocol#llama.cpp backend#LM Studio integration

An on-device document search engine that combines keyword and semantic search with local GGUF models, plus MCP access for AI agents. MIT licensed.

30.1KUpdated 3 weeks agoMIT

macOS#GGUF#Hugging Face integration#Hybrid search

Self-hosted AI gateway with OpenAI and Anthropic API compatibility, Ollama and vLLM support, caching, failover, and per-team usage tracking. MIT licensed.

1.2KUpdated 1 day agoMIT

macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing

Self-hosted LLM inference for Kubernetes with NVIDIA, AMD and Apple Silicon support, OpenAI-compatible APIs, and an Apache 2.0 license.

223Updated 17 hours agoApache-2.0

macOS · Linux#GGUF#Git integration#Guardrails

Self-hosted AI chat and document Q&A runs on Linux, macOS and Windows with local or cloud models. Apache 2.0 licensed; archived and no longer maintained.

12KUpdated 12 months agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access

An open source LLM programming language that combines Python logic with output constraints and supports local llama.cpp and Transformers models or cloud APIs.

4.2KUpdated 1 year agoApache-2.0

Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration

An open-source AI coding assistant for VS Code using Ollama, llama.cpp, LM Studio or hosted APIs, with an MIT-licensed self-hosted team gateway.

3.7KUpdated 1 day agoMIT

Docker · Web · VS Code#Git integration#Hybrid search#llama.cpp backend

An open source AI character interface you can run locally, with voice chat, VRM avatars, and support for Ollama, llama.cpp, and cloud APIs.

1.6KUpdated 1 year agoMIT

Windows · Docker · Web#llama.cpp backend#LM Studio integration#Multimodal input

A browser-based LLM frontend for llama.cpp, koboldcpp, AI Horde and OpenAI-compatible APIs, with offline use and an AGPL-3.0 license.

752Updated 9 months agoAGPL-3.0

Web#llama.cpp backend#OpenAI-compatible API#Persistent memory

A native macOS AI chat client that connects to Ollama, LM Studio and cloud APIs, with inline writing in other apps and locally stored data.

mindmac.appDesktop Chat Apps

macOS#llama.cpp backend#LM Studio integration#MLX

Desktop AI app powered by llama.cpp, with downloads for Windows and Intel or Apple Silicon Macs. The desktop app is deprecated and unsupported.

desktop.backyard.aiDesktop Chat Apps

macOS · Windows#llama.cpp backend

A self-hosted LLM chat interface that runs models through llama.cpp in Docker without remote API keys. Source is licensed under MIT and Apache 2.0.

5.7KUpdated 1 year agoApache-2.0

Windows · Docker · Web#llama.cpp backend

A self-hosted ChatGPT alternative that runs Llama 2 and Code Llama locally, with an MIT license and an OpenAI-compatible API.

10.9KUpdated 3 years agoMIT

macOS · Docker · Web#GGUF#llama.cpp backend#OpenAI-compatible API

An offline AI assistant for Android and iOS that runs models on your phone, keeps chats encrypted on-device, and offers optional hosted models.

layla-network.aiAI Characters and Roleplay

iOS · Android#Code execution#GGUF#llama.cpp backend

Local LLM web interface for Windows, macOS and Linux. Use GGUF models, Ollama or cloud APIs, with local chat storage and an Apache 2.0 license.

4.8KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

An on-device AI SDK that runs text, image and audio models on macOS, Windows and Linux, with GGUF, MLX and an OpenAI-compatible API.

qualcomm/GenieXInference Libraries and Bindings

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

Self-hosted AI data assistant that queries databases, analyzes files, and generates reports with local models or cloud APIs. Open source under MIT.

20.1KUpdated 2 days agoMIT

macOS · Linux · Docker · Web#Code execution#llama.cpp backend#OpenAI-compatible API

A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

2.1KUpdated 3 weeks agoApache-2.0

Linux#GGUF#Guardrails#Hugging Face integration

An open source LLM evaluation framework for local models and hosted APIs, with GGUF, Hugging Face transformers and llama.cpp support. MIT licensed.

14.1KUpdated 2 weeks agoMIT

macOS#Batch processing#GGUF#Hugging Face integration

A local LLM app that runs models offline on iOS and macOS, supports text and vision models, and uses ggml and llama.cpp under the MIT license.

2.1KUpdated 8 months agoMIT

macOS · iOS#llama.cpp backend#Multimodal input#RAG

A local LLM family for chat, coding and multilingual tasks, with GGUF and Hugging Face formats for CPU or GPU use and support for llama.cpp and MLX.

127Updated 12 months ago

macOS#GGUF#Hugging Face integration#llama.cpp backend

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

A bilingual local LLM family for English and Chinese, with chat and base models, version-specific licensing, and quantized variants for consumer GPUs.

7.8KUpdated 2 years agoApache-2.0

Docker#Hugging Face integration#llama.cpp backend#Multilingual

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

8.9KUpdated 8 months agoApache-2.0

Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration

Open-source text-to-speech built on Llama, with local inference, voice cloning and streaming audio. Uses Apache 2.0; Baseten offers cloud hosting.

6.3KUpdated 10 months agoApache-2.0

#Hugging Face integration#llama.cpp backend#LoRA

An open-source terminal AI agent that runs on your machine with Ollama or cloud models, edits files, executes code, and supports persistent tasks.

4.4KUpdated 20 hours agoMIT

macOS · Windows · Linux · Web · JetBrains#Agent Client Protocol#Code execution#Git integration

Self-hosted AI chat interface under Apache 2.0. Connect it to Ollama, llama.cpp or cloud APIs, with optional model routing and MCP tools.

11KUpdated 1 day agoApache-2.0

Docker · Web#llama.cpp backend#MCP#Multi-user access

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

9.4KUpdated 3 weeks agoMIT

Docker#Batch processing#GGUF#Hugging Face integration

An open-source local LLM stack manager that connects Ollama, llama.cpp and AI apps through Docker Compose. Includes a CLI and companion app.

3.2KUpdated 5 days agoApache-2.0

macOS · Linux · Docker#GGUF#llama.cpp backend#MCP