Ollama JavaScript connects Node.js and browser applications to models running through Ollama. It's for developers building chat interfaces, AI agents or other apps that need a local LLM backend. The library is open source under the MIT license, with TypeScript types and an API that follows Ollama's REST interface.
Chat and text generation support streamed responses, so apps can display answers as they arrive. Developers can include images in prompts, request JSON output and use model tool calls. Thinking controls depend on the chosen model. The library also generates embeddings for search and retrieval features, and can cancel active response streams.
Model management is included: apps can download, inspect, copy and delete models, or create derivatives with custom prompts and LoRA adapters. A configurable server address lets the same client connect to Ollama on the user's machine or another server.
Local requests go to the chosen Ollama server. Cloud models, such as gpt-oss:120b, run on Ollama's infrastructure, either through a local Ollama connection or directly through its hosted API. Cloud access requires sign-in or an API key. The client also provides web search, which requires an Ollama account, and web page fetching.
For classification and scoring, System One works with compatible local models such as nimble. It returns structured answers for choice, probability and score questions in a single JSON response; it doesn't support cloud models or streaming.
Claim this page and we'll verify you by hand. Ollama JavaScript gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Ollama JavaScript?Promote it
Something wrong or outdated on this page?
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
4.2KUpdated 1 year agoApache-2.0
Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration
19.2KUpdated 2 weeks agoApache-2.0
Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.
WebLLM runs language models directly in a user's browser, using WebGPU for GPU acceleration. It's an open-source engine for developers building web-based AI assistants and Chrome extensions that process prompts on the user's device rather than an inference server. The project uses the Apache 2.0 license.
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.