
Transformers is a Python library for developers and researchers who want to run pretrained AI models or train their own on hardware they control. It covers language, images, audio, video and multimodal work through a shared way of defining models. The library runs in a local Python environment; pretrained checkpoints are available from the separate Hugging Face Hub.
For inference, its Pipeline interface handles tasks such as text generation, speech recognition, image segmentation and questions about documents. Its generation tools support streaming output from large language and vision language models. Training is covered too: Trainer works with PyTorch models and supports distributed training, mixed precision and FlashAttention.
A model definition in Transformers can also be used by tools across the wider ecosystem, including Axolotl and Unsloth for training, vLLM and SGLang for inference, and libraries such as llama.cpp and mlx. That makes it useful if you expect to experiment with a model before choosing how to train or serve it. Researchers can inspect and customize model code, while developers can start from pretrained checkpoints instead of building a model from scratch. Transformers is released under Apache 2.0.
Claim this page and we'll verify you by hand. Transformers gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Transformers?Promote it
Something wrong or outdated on this page?
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
14.2KUpdated 5 days ago
#Hugging Face integration#Multimodal input
OpenCLIP is a Python and PyTorch library for developers and researchers who want to match images with text on their own hardware. It implements OpenAI's CLIP approach: images and descriptions become numerical representations that the model can compare. This supports image search and zero-shot classification, where text labels define the categories without a separate classifier trained for each task.
1.6KUpdated 9 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.