Tools tagged with "Streaming inference"

50 tools
Favicon of LangChain

LangChain

5 videos
An MIT-licensed Python framework for AI agents and LLM apps, with model, tool and data integrations. LangChain.js serves JavaScript and TypeScript.

147.3KUpdated 1 day agoMIT

#Human approval#RAG#Streaming inference

JavaScript client for Ollama in Node.js and browsers. Connect to local models or Ollama's cloud with an open-source, MIT-licensed library.

4.4KUpdated 2 days agoMIT

Web#LoRA#Multimodal input#Ollama integration

A mobile AI assistant that runs GGUF models on iOS and Android. Core chat works offline after a model download and needs no account.

8.5KUpdated 2 days agoMIT

iOS · Android#GGUF#Hugging Face integration#llama.cpp backend

Self-hosted speech API for transcription, translation and speech generation. Runs via Docker on CPU or GPU with faster-whisper, Kokoro and Piper.

3.7KUpdated 5 months agoMIT

Docker#OpenAI-compatible API#Streaming inference

Favicon of llama-swap

llama-swap

1 video
A local AI proxy that switches models on demand through OpenAI and Anthropic compatible APIs. Runs on macOS, Windows, Linux and FreeBSD under MIT.

5.8KUpdated 2 days agoMIT

macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend

An open-source Python framework for AI chat apps you can self-host, with customizable interfaces, OAuth and integrations including LangGraph and HuggingFace.

12.5KUpdated 1 month agoApache-2.0

Web#LLM tracing#Multi-user access#Streaming inference

Favicon of MLX LM

MLX LM

2 videos
A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

7.2KUpdated 1 day agoMIT

macOS#Batch processing#Distributed execution#Hugging Face integration

.NET library for local LLM apps using Ollama, with streaming chat, embeddings and model management. Open source under MIT; also supports Ollama cloud.

1.4KUpdated 2 months agoMIT

#Multimodal input#Ollama integration#Streaming inference

An MIT-licensed LLM data extraction library that validates structured outputs with Pydantic and works with Ollama, llama-cpp-python, vLLM and cloud APIs.

14KUpdated 3 weeks agoMIT

#llama.cpp backend#Ollama integration#Streaming inference

Open-source text-to-speech toolkit for local speech generation, voice cloning and model training on Linux, macOS and Windows, licensed under MPL-2.0.

2.3KUpdated 4 months agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning

Favicon of WebLLM

WebLLM

2 videos
A local LLM engine that runs models in the browser with WebGPU acceleration, OpenAI API compatibility, and an Apache 2.0 license.

19.2KUpdated 2 weeks agoApache-2.0

Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output

An open-source Python AI framework for self-hosted agents and document search, with local model support and an Apache 2.0 license.

26.6KUpdated 1 day agoApache-2.0

Docker#Guardrails#Hugging Face integration#Hybrid search

Python client for Ollama with local and cloud model access, streaming chat, embeddings, and async support. Open source under the MIT license.

10.6KUpdated 2 days agoMIT

#Batch processing#Ollama integration#Streaming inference

A free TypeScript library for AI apps and agents, with streaming, provider switching, and chat UI hooks for React, Next.js, Svelte, and Vue.

27.1KUpdated 3 hours ago

#Agent Skills#Streaming inference#Structured output