LM Studio: Intel Arc A770 and RX 6800 dual-GPU setup

Learn how the speaker splits Qwen models across AMD and Intel GPUs in LM Studio, reporting 50 tokens per second for a 28 GB Qwen3.5 model.

Player not loading? Watch on YouTube

The speaker pairs a secondhand Intel Arc A770 bought for $200 with an AMD RX 6800 to get 32 GB of combined VRAM. This LM Studio tutorial compares each card alone with a mixed GPU setup for a local LLM that exceeds either card's 16 GB capacity.

The Windows system uses a B550 motherboard, Ryzen 7 3700X and 32 GB of DDR4-3200 memory. Its second slot runs at PCIe 3.0 x4. The speaker recommends checking card clearance, enabling Above 4G Decoding and Resizable BAR in BIOS, and checking PCI Express Link State Power Management if an idle secondary GPU becomes unavailable.

For Qwen3 14B Q4_K_M at 4K context, the reported speeds are about 40 tokens per second on the RX 6800 and 25 on the A770 through Vulkan. Converting the model to OpenVINO raises the A770 result to just under 30. These are results from the speaker's machine; the cited Nvidia figures come from different systems.

The dual-GPU configuration uses Vulkan, an even split and GPU KV cache offload. The speaker disables flash attention because it produces garbled output with the A770 in LM Studio 0.4.9 and the Intel drivers used. Qwen3 32B Q5_K_M, about 23 GB, reaches 13 tokens per second. Qwen3.5 35B MoE Q5_K_M, about 28 GB, reaches 50 after GPU offload is set to maximum. The speaker checks one benchmark answer, so this is not a broad assessment of model quality.