
Hunyuan-A13B is Tencent's downloadable language model for developers and researchers building self-hosted AI applications. It supports general text tasks, reasoning and agent workloads, with a choice between quick responses and more deliberate reasoning. It's part of the broader Hunyuan model family, which Tencent also offers through its website.
The model uses a Mixture-of-Experts architecture: each token activates only part of the model rather than all its parameters. That reduces computation during inference, though the full model still contains the weights of its other experts. This approach matters for teams comparing the resources needed to serve a large language model on their own infrastructure.
A 256K context window lets it work with long documents and substantial task context. Its fast and slow thinking modes give applications control over how much reasoning the model performs. Agent capabilities are another focus, with evaluations covering function calling and task execution.
Hunyuan-A13B works with Transformers and supports deployment through vLLM. Tencent provides FP8 and INT4 quantized models and documents INT8 deployment, giving users alternatives to the full-precision weights that reduce memory demands. The broader Hunyuan offering includes Hy3 for language tasks and Hy Vision for image understanding; those are separate models from the downloadable A13B model described here.
The weights use the Tencent Hunyuan-A13B Community License, with conditions separate from deployment-library licenses.
Claim this page with an email at hunyuan.tencent.com. Hunyuan LLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Hunyuan LLM?Promote it
Something wrong or outdated on this page?
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
20.4KUpdated 2 months agoApache-2.0
macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.