Player not loading? Watch on YouTube
Autodesk researchers explain how they use Ray across a 3D generation pipeline, from geometry tokenization to distributed evaluation and deployment. The session focuses on AWS custom accelerators for training and inference rather than a desktop local LLM setup.
For Qwen3 fine-tuning, they test 8B and 32B models on Trainium 1 with Ray Train and Hugging Face Optimum Neuron. Their instance has 32 NeuronCores and 512 GB of device memory. Tensor parallelism of 32 lets them fit the 32B model on one instance. They report that longer sequences improve utilization within the memory budget, while scaling from one to four nodes reduces job time by 3.47 times and raises cost per million tokens by 15% in their experiment.
The GPU comparison has limits: the speaker says the configurations favor Trainium. H100 leads in tokens per second per device, and Trainium adapters perform worse in their final evaluation for an unresolved reason. Compilation startup also makes short training runs less economical.
The inference section covers kernel layout, static shapes and deployment through Ray Serve after Neuron compilation. In a Qwen3 4B benchmark, Inferentia2 delivers comparable time to first token at 80% of the GPU instance cost at concurrency 200, but trails on most other metrics. The speaker recommends testing fixed-input workloads such as diffusion-based 3D generation and notes difficulties with KV-cache decoding and dynamic shapes.