Player not loading? Watch on YouTube
This tutorial sets up RWKV 7 as a local LLM through Ollama, then connects it to SillyTavern for character chat. The speaker demonstrates the 13.3B model on a 16 GB Mac Mini and reports a download of roughly 8.4 GB. Ollama must already be installed; the video refers viewers elsewhere for that setup.
The walkthrough uses the heredos/rwkv7 model listing, which includes 2.9B, 7.2B and 13.3B variants. The speaker appends the 13.3b tag to select the larger model and says the untagged command defaults to 2.9B. After a terminal check and a conversation in Ollama, the tutorial opens SillyTavern's connection profile, selects Text Completion with the Ollama API type, and chooses RWKV.
The speaker describes Ollama's context limit as one million tokens, but estimates that SillyTavern's context control stops around 524,000 unless unlocked. Those figures are presented as limits in this setup. The demonstration does not establish perfect recall across a million-token conversation.
The second half explains RWKV's recurrent state and sequential token processing. The speaker attributes its long-context behavior to linear processing cost and constant memory use, while acknowledging that the stored state holds finite information. Practical caveats include odd responses to sparse prompts, limited role-play quality in the base model, and the speaker's advice against using these small variants in an AI agent yet.