llama.cpp, AnythingLLM, Pi and n8n: local stack setup

Learn to connect chat, document search, coding and email automation to llama.cpp, with a port 8080 example and an hourly Gmail labeling workflow.

Player not loading? Watch on YouTube

This tutorial connects a self-hosted AI stack to one OpenAI-compatible REST endpoint. llama.cpp runs the models, while llama-server handles serving and routing. The speaker describes its built-in router as experimental and presents llama-swap as an alternative configured through a single file.

For chat, the walkthrough deploys AnythingLLM with Docker Compose and selects its generic OpenAI provider. The example points to the server on port 8080, then sets context and output token limits. Document search uses the default LanceDB vector database and embedder. After uploading research papers, the speaker switches the workspace from agent to chat mode and demonstrates an answer with a source document reference.

The coding assistant setup installs Pi through npm, adds a llama.cpp plugin and configures the server URL with /v1. The /models command selects a served model. The speaker reports that Pi analyzed older codebases and fixed configuration errors in an Angular project; these are examples from his use.

The n8n workflow checks email hourly and lets an AI agent label relevant Gmail messages. It uses the local endpoint through OpenAI credentials with the Responses API disabled. The speaker says this avoids cloud AI processing, though the workflow still connects to Gmail. Closing advice covers a dedicated machine, BIOS recovery after power loss, container management and Tailscale remote access.