3.7KUpdated 11 months agoApache-2.0
Web#Hugging Face integration#LoRA
Mochi 1 is a text-to-video model for creators and developers who want to generate videos on their own hardware or adapt a model to their own footage. Genmo releases it under Apache 2.0, with downloadable weights and code for local use. Genmo also offers a hosted playground for trying the model in a browser.
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
1.1KUpdated 2 years agoApache-2.0
Web#Batch processing#Hugging Face integration#LoRA
CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.
huggingface.coImage Generation Models
Chroma is an 8.9B text-to-image model based on FLUX.1-schnell for creators and developers who want local image generation. The original model repository is deprecated; its maintainer recommends Chroma1-HD, Chroma1-Base or Chroma1-Flash instead. Those are separate successor checkpoints.
27.3KUpdated 9 months agoMIT
Web#Hugging Face integration#Image-to-image
Stable Diffusion XL is an image generation model for creators and developers who want to generate images on their own hardware or servers. It supports text-to-image generation and image-to-image sampling, so you can start with a written prompt or an existing image.
8.4KUpdated 8 months agoApache-2.0
Web#Image-to-image#LoRA#Multimodal input
Qwen-Image is an open-source image generation and editing model you can deploy locally. It's for developers and creators who want to generate images from text or revise existing pictures on their own hardware. Its text rendering capabilities, especially for Chinese, make it relevant for images that need readable lettering alongside visual content.
73.5KUpdated 4 years ago
#Guardrails#Hugging Face integration#Image-to-image
Stable Diffusion 1.5 is an AI image generation model for creators and developers who want to generate images on their own hardware. It turns text prompts into images and supports text-guided changes to existing pictures, including turning rough sketches into detailed artwork. Local inference keeps that image generation work on your machine.
huggingface.coOCR and Document Scanning
Linux#Batch processing#Hugging Face integration#Multimodal input
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multimodal input#Tool calling
Pixtral is a family of Mistral models for developers who want to run multimodal AI on their own hardware. The associated mistral-inference project is archived and no longer maintained. That status applies to the inference library.
3.9KUpdated 1 week agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
SmolVLM is a compact vision language model from Hugging Face for developers building local AI applications that work with images and text. It can describe pictures, answer questions about diagrams, and read information from documents such as invoices. Its small memory footprint makes on-device use practical on laptops and smaller local setups.
huggingface.coCoding Models
#Hugging Face integration
Qwen2.5-Coder is a family of code-focused language models for developers who want to run a coding assistant on their own hardware or build a self-hosted code agent. It covers code generation, reasoning about code and fixing errors. The 32B instruction-tuned model is open source under Apache 2.0 and accepts requests in a chat format.
huggingface.coOpen-Weight LLMs
#Hugging Face integration#LoRA#Multilingual
Jamba is AI21's language model family for teams building AI applications on their own servers. The documented Large 1.7 model combines Mamba state-space models with Transformer attention to process long context efficiently. A 256K-token context window makes it relevant for work that depends on lengthy documents, such as investment research, due diligence and reviewing procurement responses.
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Tool calling
Codestral and Devstral are Mistral coding models with downloadable weights for local or self-hosted development tools. Codestral focuses on code generation and fill-in-the-middle completion, where the model fills a gap between existing code. Devstral is designed for software engineering agents that explore a repository, use tools and edit multiple files.
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multimodal input#Tool calling
Mistral Small and Large are downloadable language models for developers building chat, reasoning and tool-using applications on their own infrastructure. Capabilities and hardware requirements depend on the release. Mistral Small 3.1 adds image understanding to text generation, while Mistral Large 2 is a larger text model.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
26.5KUpdated 3 weeks agoApache-2.0
macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend
MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
11.1KUpdated 11 months ago
#Hugging Face integration#Quantization#Tool calling
Kimi K2 is Moonshot AI's language model series for developers building coding assistants and AI agents, and researchers who want a foundation model to customize. You can run its checkpoints on your own infrastructure or use Moonshot's hosted API. Local inference runs on your hardware; the hosted API sends requests to Moonshot's service.
3.5KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multimodal input
PaliGemma is a family of downloadable vision-language models for developers and researchers building applications that work with images and text. It combines SigLIP's image processing with Gemma's language capabilities to answer questions about visual content. Its main appeal is task-specific fine-tuning: you can adapt a base model to your own image data and intended use.
1.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
PixArt-Sigma is a text-to-image diffusion model that supports generation at 2K and 4K resolutions on your own machine or server. It's aimed at developers and researchers who want pretrained models they can run themselves, along with code for training and adapting them. The Python code is open source under Apache 2.0.
25KUpdated 2 years agoApache-2.0
macOS · Web#LoRA#Multimodal input#Quantization
LLaVA is a family of vision-language models for researchers and developers who want to ask questions about images on their own hardware. It pairs a CLIP vision encoder with a language model to support image descriptions, visual reasoning and reading text in pictures. Its Python code is open source under Apache 2.0; the project places research-use restrictions on its data and checkpoints, with additional terms from the underlying models.
820Updated 1 year ago
#Quantization#Tool calling
Hunyuan-A13B is Tencent's downloadable language model for developers and researchers building self-hosted AI applications. It supports general text tasks, reasoning and agent workloads, with a choice between quick responses and more deliberate reasoning. It's part of the broader Hunyuan model family, which Tencent also offers through its website.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
2.1KUpdated 3 years agoApache-2.0
#Hugging Face integration#LoRA#Quantization
StarCoder2 is a family of code generation models for developers who want to run code completion on their own hardware or adapt a model to their code. It predicts code continuations rather than following conversational instructions, so it's suited to completion workflows rather than a chat-based coding assistant.
7KUpdated 11 months agoMIT
#Hugging Face integration#Multilingual
DeepSeek-Coder is a family of code language models for developers who want downloadable models for their own coding tools. It generates code, completes existing code, and fills gaps within a file using the surrounding context. DeepSeek also offers a hosted chat interface; the downloadable weights are separate from that service.
3.9KUpdated 1 month agoMIT
#Hugging Face integration#Ollama integration#Works offline
Phi-4 is Microsoft's small language model for complex reasoning and math problem solving. It's for developers building generative AI applications who want to run a model on their own hardware, rather than depend on a cloud connection. The original Phi-4 has 14 billion parameters and uses the MIT license; hardware needs depend on the chosen model and precision.
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
7.8KUpdated 2 years agoApache-2.0
Docker#Hugging Face integration#llama.cpp backend#Multilingual
Yi is a family of open-weight language models from 01.AI for developers, researchers, and businesses that want to run English and Chinese models on their own hardware. It includes models for conversation and base models for fine-tuning, with license terms that depend on the release. Yi-1.5 code and weights use Apache 2.0; the original Yi releases have separate community terms.
3.2KUpdated 1 year agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
MiniMax-M1 is an open-source reasoning model for developers building agents or working on complex software and mathematical problems. Its million-token context window makes it a candidate for tasks with long inputs that also need extended reasoning. You can serve the model on your own infrastructure through vLLM or use it through Transformers.
7.3KUpdated 1 year agoApache-2.0
#Hugging Face integration#Tool calling#Web search
InternLM is a family of downloadable language models for developers and researchers building their own AI applications. It includes models for general conversation, complex reasoning and coding, with separate base and chat variants for customization or use in an assistant.
9.2KUpdated 2 weeks agoApache-2.0
#ControlNet#LoRA#Multimodal input
Sana is an open-source framework for running image and video generation on your own hardware, with image models small enough for laptop GPUs. It's aimed at creators who want local AI generation and developers who need training and inference pipelines for their own models. The code uses the Apache 2.0 license.