Favicon of OpenVINO

OpenVINO

Open-source AI inference toolkit for local or self-hosted deployment on Linux, Windows and macOS, with CPU, Intel GPU and NPU support.

Screenshot of OpenVINO website

OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.

The toolkit covers local LLM inference, image generation, speech recognition and computer vision. OpenVINO GenAI handles generative models, while the base package handles conventional AI workloads. OpenVINO Physical AI supports vision-language-action models for robots.

Models can come from PyTorch, TensorFlow, ONNX, Keras, PaddlePaddle or JAX/Flax. Optimum Intel connects it to Hugging Face transformers and diffusers, and pre-optimized OpenVINO models are available on Hugging Face. Converted models can run without their original training frameworks. PyTorch applications can also use OpenVINO through torch.compile().

For deployment, applications can run inference through the local runtime or send requests to OpenVINO Model Server on a separate server, including Kubernetes and cloud deployments. Python, C++ and C APIs support integration into existing applications.

Model compression reduces resource demands. Compiled-model caching helps shorten startup, and the runtime can begin inference on the CPU while another device prepares the model. OpenVINO collects performance and usage telemetry, which users can opt out of.

Similar to OpenVINO