Slotstream, Codewhale and Utopia: local AI overview

Explore local model runners and agent tools, including Slotstream's SSD streaming on Mac and a 27B model recipe for a 24 GB RTX 3090.

Player not loading? Watch on YouTube

This GitHub Trending Weekly roundup introduces 34 projects, with several entries relevant to running AI on your own hardware. It is a brief overview rather than a setup tutorial, and the performance figures are reports from the speaker or project teams.

Slotstream streams 104 GB of model weights from SSD on a Mac that cannot hold the whole model in RAM. The speaker reports about 12 tokens per second for warm decoding on a 48 GB M5 Pro Mac and a 16 GB RAM floor. A separate qwen38-27b-rtx3090 recipe targets a 27-billion-parameter local LLM on a 24 GB GPU. Its reported 417 tokens per second is batched throughput; the segment describes reclaiming memory by quantizing both embedding matrices to int8.

Codewhale is a Rust terminal coding assistant that can use hosted APIs or local backends. The overview covers switching models during a session, approval modes and workspace recovery. Utopia is a self-hosted RAG platform with historical fact tracking, hybrid search and source citations. The speaker says uncertain extractions enter a review queue.

The roundup also covers SkillRadar's offline static analysis of skills and Skill Cabinet's installed-skill browser. Skill Cabinet deletes folders from disk without undo. CC prune removes processed tool output from Claude transcripts, with snapshots for rollback.