Player not loading? Watch on YouTube
The speaker explains why GPT-Oss:120b remains the model for his public tech support bot on a Strix Halo machine with 128GB of unified memory. His local LLM choice depends on response speed, instruction following and reliable tool calls rather than model age.
With llama.cpp, he reports 688 tokens per second for prefill and 53 for decode. His Qwen-3.8-Flash-Next Q2 GGUF test used 75GB of memory and reached 311 prefill and 31 output tokens per second. These are results from his setup, not general benchmarks. He did not establish whether Q2 gave better answers.
The machine reserves 96GB for VRAM and 32GB for RAM. Flash Next Q4 has about 104GB of weights, before KV cache. He tried GTT memory mapping to accommodate larger models, but GPT-Oss:120b dropped to 103 prefill and 15 decode. That result led him to abandon the Q4 test.
For knowledge updates, his AI agent uses web search and a RAG database. He describes correcting support answers through retrieval data while keeping the model fixed. External sources do not guarantee accurate answers; he also notes that models can ignore retrieval instructions.
Public deployment adds security work. He describes separate screening and action steps with isolated permissions, roughly 13K of directives and about 16K of context per interaction. His replacement criteria include resistance to prompt injection and context corruption, alongside dependable tool calling.