Player not loading? Watch on YouTube
The speaker explains a self-hosted Hermes setup on an NVIDIA DGX Spark, using OpenShell to separate trusted automations from experiments. He also describes NemoClaw as the route for running OpenClaw through the same containment layer. This is a walkthrough of his security choices, rather than proof that the arrangement is universally safe.
Hermes Core and Hermes Dangerous run in separate Docker containers. Core can read and write files and execute terminal commands, but its network policy rejects most URLs by default. Dangerous adds vision and unrestricted browser access, while its policy blocks connections to Core's ports. Both containers share a local LLM served by vLLM under the hermes-default model ID. SearXNG supplies search through external providers, so local inference does not make the whole workflow offline.
The speaker tests workflows in Dangerous, then moves scheduled jobs to Core with the permissions they need. He discusses unprivileged containers, credentials outside the sandbox and state snapshots for recovery. A Tailscale dashboard provides access; Telegram is more convenient for him but introduces phone access and external server concerns.
The demonstrations also expose limits. Cloudflare blocks one browser request, and an agent misidentifies its model after a switch. On this setup, the reported benchmarks reach about 27 tokens per second with Qwen3.8-27B and 135 with Nemotron 3.5 Lightning, which the speaker says lacks vision support.