
Phi-4 is Microsoft's small language model for complex reasoning and math problem solving. It's for developers building generative AI applications who want to run a model on their own hardware, rather than depend on a cloud connection. The original Phi-4 has 14 billion parameters and uses the MIT license; hardware needs depend on the chosen model and precision.
Local deployment gives you control over the model and its environment, and it can run without cloud connectivity. Phi models are available through Ollama and Hugging Face. Microsoft also offers hosted inference through Microsoft Foundry and Azure APIs, where the model runs in the cloud rather than on your device.
You can fine-tune Phi models with domain-specific data to adapt them to an organization's needs. The family emphasizes resource efficiency and low latency for applications that need quick responses, such as real-time guidance and autonomous systems. Phi-4's particular focus is reasoning and mathematics, so it's a candidate for applications where solving problems matters more than handling media inputs.
The related models serve different needs. Phi-4-mini supports function calling and natural-language instructions, while Phi-4-multimodal accepts text, images, and audio. The multimodal model covers speech recognition, translation, and image analysis, including OCR and interpretation of charts and tables.
Claim this page with an email at azure.microsoft.com. Phi-4 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Phi-4?Promote it
Something wrong or outdated on this page?
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
20.4KUpdated 2 months agoApache-2.0
macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
27.7KUpdated 9 months ago
#Batch processing#GGUF#Hugging Face integration
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
Qwen3 is a family of language models from Alibaba Cloud’s Qwen team for people who want to run models locally or on their own servers. It spans smaller and larger dense models as well as mixture-of-experts models. The weights are publicly available.
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.