Qwen Code with local models: coding test and limitations

See Qwen Code use qwen3-coder-30b-a3b-instruct on an RTX 5090, with a Frozen Lake coding test that exposes prompting and execution problems.

Player not loading? Watch on YouTube

This first-impressions session tests Qwen Code with qwen3-coder-30b-a3b-instruct through LM Studio on an RTX 5090. The speaker reports roughly 22 GB of VRAM use for the model and says it can also fit on a 3090 or 4090. A later reading of 23.7 GB includes recording the desktop, so it is not an isolated model memory measurement.

The coding assistant works on a Frozen Lake Q-learning project. A vague request for test.py and test.sh produces a test suite with dependency problems rather than the visual demonstration the speaker intended. After clarification, the task becomes displaying the game with its trained policy. The eventual visualization runs, although the game agent falls into a hole.

The local LLM responds quickly in this session, but the speaker finds it needs more explicit prompts than Claude. It rereads files and claims to be working in a simulated environment despite having command access. The speaker suspects a context problem; the session does not establish that diagnosis. An explicit reminder to run commands prompts another testing attempt.

For people looking to run models locally, the example shows both useful code generation and the need to check execution claims. The speaker reports no token charges for this hardware-based session and sees the model as a helper for bugs and individual functions. The video ends with a separate car-racing demonstration that pairs world models with Q-learning.