Favicon of gpt-oss

gpt-oss

Open-weight local LLMs under Apache 2.0, with 20B and 120B models that work with Ollama, LM Studio and vLLM.

Screenshot of gpt-oss website

gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.

The smaller gpt-oss-20b suits local applications where response time matters. gpt-oss-120b is built for heavier general-purpose reasoning and can run on a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X. Both models are intended for desktops, laptops and data centers, though the hardware needed depends on the model and software used to run it.

The models follow instructions and can use tools, including Python code execution and web search. Developers can adjust how much reasoning they do, fine-tune them for a specific use, and inspect their full chain of thought when debugging. Web search needs an internet connection; the models can also be used without that tool.

Ollama and LM Studio provide routes to local use. The repository also includes reference implementations for NVIDIA GPUs on Linux and Apple Silicon through Metal, plus examples using vLLM and Transformers. It includes a terminal chat client and a basic Responses API compatible server for developers building their own applications. Both models use OpenAI's Harmony format for conversations and tool calls.

Similar to gpt-oss