Hermes Agent setup with vLLM, Proxmox and Open WebUI

Learn to connect Hermes Agent to local vLLM inference and Open WebUI, with an Ubuntu 24.04 VM, API configuration and timeout troubleshooting.

Player not loading? Watch on YouTube

This tutorial sets up Hermes Agent in a Proxmox virtual machine, connects it to vLLM inference, and adds an Open WebUI interface. The speaker uses Ubuntu Server 24.04 with eight CPU cores, 16 GB of assigned RAM and a 32 GB disk. These are the demonstrated VM settings, not stated minimum requirements. GPU configuration belongs to the inference server and needs adjustment for other hardware.

The installation covers OpenSSH, system updates, the QEMU guest agent and Git before starting Hermes. A custom provider endpoint uses port 9876 with the /v1 path. The speaker then reviews configuration files, diagnostic commands, updates, terminal tools and session controls. For the self-hosted browser interface, Hermes exposes an API server on 0.0.0.0:6789, which Open WebUI connects to through its admin connection settings.

The local AI agent retrieves a Hugging Face news article in the demonstration, but repeated attempts to generate a cat SVG time out. The speaker sets the terminal timeout to 120 seconds and reports slower inference and declining tool-call reliability around 64K context with Qwen 3.5 27B. Those observations describe this setup rather than a universal model limit. Backups and snapshots provide a rollback point before agent-proposed configuration patches; the speaker also reports that manual context compression restored some capability.