
Roboflow Inference is a self-hosted computer vision server for teams building camera and image analysis systems. It runs on your own computer, server, or edge device and combines model predictions with workflows for tracking, counting, measuring, and responding to events. Roboflow also offers hosted servers and a Serverless Cloud API, where processing runs on its infrastructure.
Workflows let you combine object detection, classification, and segmentation with OCR, barcode and QR reading, or template matching. You can use foundation models such as Florence-2, CLIP, and SAM2, or serve your own fine-tuned models. Applications include license plate reading, face blurring, background removal, and measuring how long an object stays in a zone. Custom code and models can extend the workflows.
For live video, it manages RTSP streams and webcams, with hardware acceleration and GPU batching. A Python SDK and REST API connect predictions to other software; workflows can also send email or Twilio notifications and call webhooks. You can record and analyze predictions alongside the video processing.
It supports Docker deployments on Linux, Windows, and macOS, plus NVIDIA Jetson and Raspberry Pi devices. NVIDIA CUDA provides GPU acceleration where available. Pre-trained and foundation models, public workflows, and video stream management don't require an API key. A Roboflow API key provides access to fine-tuned models, private workflows, and hosted compute; an account also enables remote stream management through the Roboflow UI. Self-hosted commercial model licensing is a separate enterprise add-on.
Claim this page with an email at inference.roboflow.com. Roboflow Inference gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Roboflow Inference?Promote it
Something wrong or outdated on this page?
982Updated 1 year ago
macOS · Windows · Linux · Docker#Home Assistant integration#Image-to-image#Multimodal input
CodeProject.AI Server gives developers a shared API for AI tasks that run on their own hardware. It's a self-hosted service for adding image analysis, text processing and generation to applications. Processing stays on the machine running the server, without cloud calls or sending data outside your device or network.
818Updated 4 years agoApache-2.0
macOS · Windows · Linux · Docker#Home Assistant integration#Works offline
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
DeepStack is a self-hosted computer vision API for developers adding image analysis to camera systems, home automation or other applications. It runs prebuilt and custom models on your own hardware and works fully offline. Image processing stays on the device or server where you host it, with no cloud service required.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.