
Safetensors is a file format and library for developers who store, share, or load AI model weights on their own hardware or servers. It avoids the arbitrary code execution risk of PyTorch's pickle-based files while supporting fast access to tensor data. The project is open source under Apache 2.0, with a Rust implementation and Python support.
Its focus is tensor storage. You can load individual tensors or slices without reading every weight into memory, which is useful when splitting a model across multiple GPUs. The format supports bfloat16 and FP8 data and doesn't impose a file size limit.
CPU loading can use cached file data without an extra copy of the tensor contents. GPU loading still requires a copy, but it can avoid holding all tensors in CPU memory at once. These distinctions matter for developers comparing formats for large models: the benefit depends on where the weights are loaded and whether the file is already cached.
Safetensors fits into existing local AI software. Projects using it include Transformers, MLX, Diffusers, ComfyUI, and text-generation-webui, so it serves both language model and image generation workflows.
The format stores tensor names, shapes, and data types in a readable JSON header alongside the tensor data. It limits header size and checks that tensor data regions don't overlap to reduce risks from malicious files. It also supports string metadata, while requiring tensors to be packed before saving.
Claim this page and we'll verify you by hand. safetensors gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find safetensors?Promote it
Something wrong or outdated on this page?
2.2KUpdated 6 days agoApache-2.0
#Ollama integration#OpenAI-compatible API#Tool calling
any-llm is a Python library for developers who want the same application to work with local LLM servers and cloud providers. It connects to Ollama and custom OpenAI-compatible endpoints, alongside OpenAI, Anthropic, Mistral and Azure / Microsoft Foundry. A shared interface reduces the provider-specific code needed to try another model or change where inference runs.
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
12.5KUpdated 1 month agoApache-2.0
Web#LLM tracing#Multi-user access#Streaming inference
1.1KUpdated 3 weeks agoMPL-2.0
Linux#Ollama integration#RAG#Semantic search
30.4KUpdated 17 hours agoMIT
#MCP#Tool calling
Composio connects AI assistants and custom agents to apps such as Gmail, Slack, GitHub, and Linear. It's for people who want their assistant to act on requests across apps, and developers who don't want to maintain each integration themselves. Its CLI gives coding agents a local interface; the standard setup uses Composio's hosted authentication and execution service. It requires an account and internet access.
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
Chainlit is an open-source Python framework for developers building conversational AI apps with their own application logic. You can run the app on your own server and give users a browser chat interface. It's a framework for creating an app, rather than a ready-made chatbot with a fixed model.
chromem-go is a vector database that runs inside your Go application, so developers can add semantic search or retrieval augmented generation (RAG) without maintaining a separate database server. It stores text alongside embeddings and retrieves related documents for use in LLM answers. Its focus is ordinary application workloads rather than collections containing millions of documents.