Tools tagged with "Prompt caching"

13 tools
Favicon of Slotstream

Slotstream

1 video
Local LLM runner for Apple Silicon Macs that runs Qwen3.8-Flash-Next from SSD. Works offline after download and connects to coding agents and chat apps.

407Updated 2 days agoMIT

macOS#MLX#Multimodal input#OpenAI-compatible API

Favicon of vLLM

vLLM

9 videos
An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

93KUpdated 2 hours agoApache-2.0

macOS · Docker#Batch processing#Distributed execution#GGUF

Favicon of llama.cpp

llama.cpp

11 videos
An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

130KUpdated 1 hour agoMIT

Web#Code execution#GGUF#Hugging Face integration

Favicon of SGLang

SGLang

3 videos
An open-source inference framework for serving language and multimodal models on your own hardware, with an OpenAI-compatible API.

36.7KUpdated 2 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

An LLM chat frontend that connects to ChatGPT, Gemini, Claude, and other models through your own API keys.

typingmind.comChat and Assistants

#Agent Skills#MCP#Multimodal input

Self-hosted LLM inference engine for Hugging Face models, with OpenAI-compatible APIs, multimodal support, and CPU or GPU execution under AGPL-3.0.

1.9KUpdated 3 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM serving stack built on vLLM, with an OpenAI-compatible API, request routing and GPU cluster monitoring. Licensed under Apache 2.0.

2.6KUpdated 20 hours agoApache-2.0

Web#OpenAI-compatible API#Prompt caching

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 21 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 24 hours agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 3 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

14.7KUpdated 21 hours ago

Docker#Batch processing#Distributed execution#LoRA

An open-source local LLM framework that splits work across CPUs and GPUs, with SGLang serving and LlamaFactory fine-tuning under Apache 2.0.

19.5KUpdated 1 week agoApache-2.0

Docker#LoRA#Multimodal input#Prompt caching

Favicon of MLX LM

MLX LM

2 videos
A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

7.2KUpdated 1 day agoMIT

macOS#Batch processing#Distributed execution#Hugging Face integration