Favicon of Qwen3

Qwen3

A language model family with public weights for local CPU or GPU use through Ollama, llama.cpp, and LM Studio, plus server deployment.

Screenshot of Qwen3 website

Qwen3 is a family of language models from Alibaba Cloud’s Qwen team for people who want to run models locally or on their own servers. It spans smaller and larger dense models as well as mixture-of-experts models. The weights are publicly available.

Qwen3 includes Instruct models for chat and writing, and Thinking models for tasks that call for deeper reasoning. The models handle coding, mathematics, science questions, instruction following, and multilingual text. General Qwen3 models can switch between thinking and direct responses, while the dedicated Instruct and Thinking variants have separate roles. That gives users a concrete choice when selecting a model for everyday conversation or more demanding problem solving.

Local use is supported on CPU or GPU through llama.cpp, Ollama, and LM Studio. Transformers is another way to run the models, while vLLM and SGLang support server deployments. Qwen3 also supports quantization with GGUF, GPTQ, and AWQ, which can make larger models practical on more hardware. Checkpoints are distributed through Hugging Face and ModelScope. Qwen Chat offers a separate online way to use the models.

Similar to Qwen3