
SmolLM3 is a 3B parameter language model from Hugging Face for developers and researchers who want to run an LLM on their own hardware. It comes as a base model and an instruction-tuned model for chat, reasoning and tool calling. Both run locally.
The instruction model has two response modes. Direct mode answers without a visible reasoning trace, while thinking mode can work through harder prompts before replying. It supports English, French, Spanish, German, Italian and Portuguese. Its 128k token context window gives it room for long documents and code, and its chat format lets applications provide tools the model can call.
SmolLM3 runs through Hugging Face Transformers or vLLM. The project shows inference on a CPU or CUDA GPU. Its repository is licensed under Apache 2.0, and Hugging Face publishes the model weights, architecture, training data mixture and training configurations. Those materials give researchers a way to examine how this compact model was built. Its architecture uses grouped query attention to reduce the memory needed for past tokens, while NoPE and YaRN support its long context window.
Claim this page and we'll verify you by hand. SmolLM3 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find SmolLM3?Promote it
Something wrong or outdated on this page?
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
3.2KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.