Favicon of Hunyuan LLM

Hunyuan LLM

A self-hosted language model for reasoning and agent tasks, with fast and slow thinking modes, a 256K context window, and Transformers and vLLM support.

Screenshot of Hunyuan LLM website

Hunyuan-A13B is Tencent's downloadable language model for developers and researchers building self-hosted AI applications. It supports general text tasks, reasoning and agent workloads, with a choice between quick responses and more deliberate reasoning. It's part of the broader Hunyuan model family, which Tencent also offers through its website.

The model uses a Mixture-of-Experts architecture: each token activates only part of the model rather than all its parameters. That reduces computation during inference, though the full model still contains the weights of its other experts. This approach matters for teams comparing the resources needed to serve a large language model on their own infrastructure.

A 256K context window lets it work with long documents and substantial task context. Its fast and slow thinking modes give applications control over how much reasoning the model performs. Agent capabilities are another focus, with evaluations covering function calling and task execution.

Hunyuan-A13B works with Transformers and supports deployment through vLLM. Tencent provides FP8 and INT4 quantized models and documents INT8 deployment, giving users alternatives to the full-precision weights that reduce memory demands. The broader Hunyuan offering includes Hy3 for language tasks and Hy Vision for image understanding; those are separate models from the downloadable A13B model described here.

The weights use the Tencent Hunyuan-A13B Community License, with conditions separate from deployment-library licenses.

Similar to Hunyuan LLM