
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.
The smaller gpt-oss-20b suits local applications where response time matters. gpt-oss-120b is built for heavier general-purpose reasoning and can run on a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X. Both models are intended for desktops, laptops and data centers, though the hardware needed depends on the model and software used to run it.
The models follow instructions and can use tools, including Python code execution and web search. Developers can adjust how much reasoning they do, fine-tune them for a specific use, and inspect their full chain of thought when debugging. Web search needs an internet connection; the models can also be used without that tool.
Ollama and LM Studio provide routes to local use. The repository also includes reference implementations for NVIDIA GPUs on Linux and Apple Silicon through Metal, plus examples using vLLM and Transformers. It includes a terminal chat client and a basic Responses API compatible server for developers building their own applications. Both models use OpenAI's Harmony format for conversations and tool calls.
Claim this page with an email at openai.com. gpt-oss gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find gpt-oss?Promote it
Something wrong or outdated on this page?
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.