OllamaSharp is a C# library for developers building .NET applications around Ollama. It connects to Ollama on your own machine or a remote server and covers the full Ollama API, including model management alongside chat and embeddings. It's open source under the MIT license.
The library supports streaming responses, so an application can display an answer as the model generates it. Its chat support keeps conversation history across turns, including tool calls and their results. You can send images to vision models, request structured JSON output and use thinking mode with reasoning models.
For apps that need to manage models as well as use them, OllamaSharp can list, download, create, copy and delete models, with progress feedback during downloads. Function calling includes source generator support, and the library supports .NET Native AOT compilation.
OllamaSharp implements the chat and embedding interfaces in Microsoft.Extensions.AI. That gives developers a way to use Ollama within applications built around Microsoft's shared AI abstractions, while retaining access to Ollama-specific capabilities. It also underpins Ollama integrations in Semantic Kernel and .NET Aspire.
Where inference runs depends on the Ollama service you connect to: local models run through your own Ollama instance, while Ollama cloud models run remotely. Cloud access uses an API key.
Claim this page and we'll verify you by hand. OllamaSharp gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OllamaSharp?Promote it
Something wrong or outdated on this page?
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
1.7KUpdated 2 days agoApache-2.0
#Batch processing#Code execution#Multimodal input
38.4KUpdated 4 days agoMIT
#Code execution#MCP#Multimodal input
14KUpdated 3 weeks agoMIT
#llama.cpp backend#Ollama integration#Streaming inference
147.3KUpdated 1 day agoMIT
#Human approval#RAG#Streaming inference
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.
DSPy is a Python framework for developers building AI applications whose tasks need clear inputs, predictable output types, and measurable results. You define what a language model should produce, then compose those tasks into a larger program. It's open source under the MIT license.
Instructor is an open-source library for developers who need structured data from local LLMs or cloud models. It turns natural-language input into typed objects that applications can use, with validation and retries built into the extraction process. Its focus is data extraction.
LangChain is an MIT-licensed open-source framework for developers building AI agents and applications powered by LLMs. It provides a shared interface for models, tools and data connections, so developers can change providers or test workflows without rebuilding the whole application.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.