
Qwen3 is a family of language models from Alibaba Cloud’s Qwen team for people who want to run models locally or on their own servers. It spans smaller and larger dense models as well as mixture-of-experts models. The weights are publicly available.
Qwen3 includes Instruct models for chat and writing, and Thinking models for tasks that call for deeper reasoning. The models handle coding, mathematics, science questions, instruction following, and multilingual text. General Qwen3 models can switch between thinking and direct responses, while the dedicated Instruct and Thinking variants have separate roles. That gives users a concrete choice when selecting a model for everyday conversation or more demanding problem solving.
Local use is supported on CPU or GPU through llama.cpp, Ollama, and LM Studio. Transformers is another way to run the models, while vLLM and SGLang support server deployments. Qwen3 also supports quantization with GGUF, GPTQ, and AWQ, which can make larger models practical on more hardware. Checkpoints are distributed through Hugging Face and ModelScope. Qwen Chat offers a separate online way to use the models.
Claim this page and we'll verify you by hand. Qwen3 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Qwen3?Promote it
Something wrong or outdated on this page?
6.7KUpdated 10 months agoApache-2.0
#Tool calling
OLMo is a family of language models for researchers and developers who want to inspect, adapt or run a model on their own hardware. Weights are downloadable. Ai2 also provides the training data, code, checkpoints and reports behind the models, giving researchers material to study the full training process rather than only the finished model.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
20.4KUpdated 2 months agoApache-2.0
macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.