Favicon of SmolLM3

SmolLM3

An open source 3B local LLM that runs on CPU or CUDA GPU and supports direct or reasoning responses, six languages and a 128k context window.

Screenshot of SmolLM3 website

SmolLM3 is a 3B parameter language model from Hugging Face for developers and researchers who want to run an LLM on their own hardware. It comes as a base model and an instruction-tuned model for chat, reasoning and tool calling. Both run locally.

The instruction model has two response modes. Direct mode answers without a visible reasoning trace, while thinking mode can work through harder prompts before replying. It supports English, French, Spanish, German, Italian and Portuguese. Its 128k token context window gives it room for long documents and code, and its chat format lets applications provide tools the model can call.

SmolLM3 runs through Hugging Face Transformers or vLLM. The project shows inference on a CPU or CUDA GPU. Its repository is licensed under Apache 2.0, and Hugging Face publishes the model weights, architecture, training data mixture and training configurations. Those materials give researchers a way to examine how this compact model was built. Its architecture uses grouped query attention to reduce the memory needed for past tokens, while NoPE and YaRN support its long context window.

Similar to SmolLM3