Favicon of Ollama Python

Ollama Python

Python client for Ollama with local and cloud model access, streaming chat, embeddings, and async support. Open source under the MIT license.

Ollama Python connects Python applications to models running through Ollama on your own machine or to Ollama's cloud service. It's for developers adding local LLM features to scripts, chat applications, or other Python projects. The library is open source under the MIT license and requires a running Ollama service for local use.

The client covers chat conversations and prompt-based text generation, with streaming responses so an application can display output as it arrives. It also generates embeddings for individual texts or batches. Synchronous and asynchronous clients let it fit into both ordinary scripts and applications that handle requests concurrently.

Model management is included. Python applications can download models, inspect their details, create variants with custom system prompts, and copy or delete them. The client can also list models currently running in Ollama. Its interface follows the Ollama REST API, with typed response objects for access to returned data.

Local requests go to your Ollama service. Cloud model requests send work to Ollama's servers, either through the local service or directly through ollama.com. Cloud access requires sign-in or an API key; supported models include gpt-oss, DeepSeek V3.1, and Qwen3-Coder.

For structured decisions, System One supports choice questions, true-or-false probabilities, and scores using compatible local models such as nimble. It returns a single JSON response and doesn't support streaming or cloud models.

Similar to Ollama Python