1.8KUpdated 7 months agoApache-2.0
#Hugging Face integration
ModernBERT is a family of open-source text encoder models for developers building document search, classification and code retrieval on their own hardware. Its longer context lets it process documents and code passages that exceed the limits of older BERT models. The code and models use the Apache 2.0 license.
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
5.1KUpdated 23 hours agoApache-2.0
Windows · Docker#Agent Skills#Hugging Face integration#ONNX
TensorRT Model Optimizer, called NVIDIA Model Optimizer or ModelOpt, is a Python library for developers preparing models for local or self-hosted inference. It reduces model size and memory use and can speed up inference through compression and other optimization techniques. It's open source under Apache 2.0.
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.
2.7KUpdated 1 week agoApache-2.0
Linux · Docker#Hugging Face integration#Quantization
Intel Neural Compressor is a Python library for developers compressing AI models for deployment on their own hardware or servers. It supports local LLM work as well as other deep learning models, with particular attention to Intel CPUs, GPUs and Gaudi accelerators. It's open source under the Apache 2.0 license.
2.3KUpdated 1 year agoMIT
Linux#Batch processing#GGUF#Hugging Face integration
AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.
5.2KUpdated 4 days agoApache-2.0
Linux · Docker · Web#Hugging Face integration#LoRA#Quantization
H2O LLM Studio is a self-hosted tool for teams that want to adapt language models to their own datasets without writing training code. Its browser interface brings training experiments, evaluation, and model testing into one place. The project is open source under Apache 2.0.
8.4KUpdated 4 months agoApache-2.0
#Batch processing#Hugging Face integration#Multilingual
Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.
22.2KUpdated 1 week agoMIT
#Hugging Face integration#Multilingual
E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.
4.7KUpdated 1 year agoMIT
#Hugging Face integration#Multilingual#Voice cloning
WhisperSpeech is a local text-to-speech system that uses OpenAI Whisper as the basis for generating speech. It's for developers and speech researchers who want to work with downloadable models on their own hardware. It supports voice cloning.
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
14.7KUpdated 1 year agoApache-2.0
Windows#Hugging Face integration#Multimodal input
Sesame CSM is an open-source speech generation model for developers and researchers building voice applications on their own hardware. It uses text and audio inputs to generate speech, with support for conversational context and different speakers. It's a model component for applications that need spoken output.
4.6KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Multilingual
AutoTrain trains custom machine learning models from your own data through a no-code interface. It's for people who need to fine-tune an LLM or build a classifier without writing a training pipeline. The local AutoTrain Advanced project is no longer maintained, so it won't receive bug fixes or new features.
2.4KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Ollama integration
UpTrain is an open source LLM evaluation tool for developers who need to measure answer quality and investigate failures in their AI applications. Its self-hosted web dashboard runs on your machine through Docker, with a Python package for evaluations inside application code. The dashboard requires no coding.
3.2KUpdated 2 years agoMIT
macOS · Web#Hugging Face integration#Multimodal input
Pyramid Flow is an open-source AI video generator for people who want to create clips on their own hardware, and for researchers working on video models. It turns text prompts into video or animates a supplied image with guidance from text. The Python code uses PyTorch and carries MIT. Model licenses are separate: the SD3-derived weights use the Stability AI Community License. Check the chosen checkpoint before deployment.
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.
12.3KUpdated 2 years agoApache-2.0
Web#Hugging Face integration#LoRA#Multimodal input
AnimateDiff adds text-driven animation to personalized Stable Diffusion models without requiring separate training for each model. It's for artists and developers who want to generate motion while keeping the visual style of a chosen image model. The Python implementation runs locally and includes a Gradio browser interface.
27.3KUpdated 9 months agoMIT
Web#Hugging Face integration
Stable Video Diffusion is an image-to-video model from Stability AI that animates a supplied still image into a short clip. It is for developers and researchers working with video generation on their own hardware. The SVD model generates 14 frames; SVD-XT is fine-tuned for 25 frames at 576×1024 resolution.
29.9KUpdated 6 months agoApache-2.0
#Hugging Face integration#Multimodal input
Open-Sora is an open source AI video generation project for developers, researchers, and creators who want to run and adapt a model on their own hardware. Its model focuses on turning reference images into video, with text prompts guiding the result. It also generates video directly from text. The code and Open-Sora 2.0 weights use Apache 2.0.
7.6KUpdated 8 months ago
Web#Hugging Face integration#Multimodal input
SkyReels-V2 is a family of downloadable video models for developers and researchers who want to run generation on their own GPUs. It supports text-to-video, image-to-video and extended video generation. Its autoregressive Diffusion Forcing architecture generates longer sequences by continuing video segments.
3.7KUpdated 11 months agoApache-2.0
Web#Hugging Face integration#LoRA
Mochi 1 is a text-to-video model for creators and developers who want to generate videos on their own hardware or adapt a model to their own footage. Genmo releases it under Apache 2.0, with downloadable weights and code for local use. Genmo also offers a hosted playground for trying the model in a browser.
13KUpdated 11 months agoApache-2.0
Windows · Web#Hugging Face integration#LoRA#Multimodal input
CogVideoX is a family of downloadable video generation models for developers, researchers and creators who want to generate clips on their own hardware. It turns English text prompts into video, animates a supplied image and can continue an existing video. A local Gradio web interface provides a browser front end for generation.
1.1KUpdated 2 years agoApache-2.0
Web#Batch processing#Hugging Face integration#LoRA
CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.
4.8KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
Lollms WebUI is a local, single-user AI interface for people who want text chat and media generation in one place. It runs on Windows, macOS and Linux, with Docker support, and lets writers, developers and other users choose models and task-specific personalities. It's free and open source under Apache 2.0. The project receives minimal maintenance.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
27.3KUpdated 9 months agoMIT
Web#Hugging Face integration#Image-to-image
Stable Diffusion XL is an image generation model for creators and developers who want to generate images on their own hardware or servers. It supports text-to-image generation and image-to-image sampling, so you can start with a written prompt or an existing image.
73.5KUpdated 4 years ago
#Guardrails#Hugging Face integration#Image-to-image
Stable Diffusion 1.5 is an AI image generation model for creators and developers who want to generate images on their own hardware. It turns text prompts into images and supports text-guided changes to existing pictures, including turning rough sketches into detailed artwork. Local inference keeps that image generation work on your machine.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
huggingface.coOCR and Document Scanning
Linux#Batch processing#Hugging Face integration#Multimodal input
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multimodal input#Tool calling
Pixtral is a family of Mistral models for developers who want to run multimodal AI on their own hardware. The associated mistral-inference project is archived and no longer maintained. That status applies to the inference library.
3.9KUpdated 1 week agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
SmolVLM is a compact vision language model from Hugging Face for developers building local AI applications that work with images and text. It can describe pictures, answer questions about diagrams, and read information from documents such as invoices. Its small memory footprint makes on-device use practical on laptops and smaller local setups.
huggingface.coCoding Models
#Hugging Face integration
Qwen2.5-Coder is a family of code-focused language models for developers who want to run a coding assistant on their own hardware or build a self-hosted code agent. It covers code generation, reasoning about code and fixing errors. The 32B instruction-tuned model is open source under Apache 2.0 and accepts requests in a chat format.
huggingface.coOpen-Weight LLMs
#Hugging Face integration#LoRA#Multilingual
Jamba is AI21's language model family for teams building AI applications on their own servers. The documented Large 1.7 model combines Mamba state-space models with Transformer attention to process long context efficiently. A 256K-token context window makes it relevant for work that depends on lengthy documents, such as investment research, due diligence and reviewing procurement responses.
1KUpdated 1 week ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Tool calling
Codestral and Devstral are Mistral coding models with downloadable weights for local or self-hosted development tools. Codestral focuses on code generation and fill-in-the-middle completion, where the model fills a gap between existing code. Devstral is designed for software engineering agents that explore a repository, use tools and edit multiple files.
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.