12.6KUpdated 3 months ago
Web#Multimodal input#Quantization
HunyuanVideo is an AI video generation model for creators and developers who want to generate footage on their own hardware. Tencent provides model weights and inference code for text-to-video and image-to-video generation, alongside a hosted web experience. Local inference runs on your GPUs; the web offering runs through Tencent's service.
3.8KUpdated 1 day agoApache-2.0
#Hugging Face integration#Quantization
LLM Compressor is an open-source Python library for developers preparing models to run on their own hardware with vLLM. It reduces model size and memory requirements through quantization, and accepts local checkpoints or models from Hugging Face repositories. It's licensed under Apache 2.0.
11KUpdated 9 months agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
LTX-Video is an AI video generation model for creators building controlled animations and developers adding video tools to their own products. You can run it locally or on your own servers using publicly available weights. The LTX family also offers a managed cloud API; local deployments can run in isolated environments without a cloud dependency.
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
2.9KUpdated 6 months agoMIT
macOS · Docker#Batch processing#Hugging Face integration#Multimodal input
Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
655Updated 2 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.
25.6KUpdated 3 hours agoMIT
#Batch processing#Hugging Face integration#Quantization
faster-whisper is a Python library for people building local speech transcription into their own software. It runs OpenAI's Whisper models through CTranslate2, with faster processing and lower memory use than the original Whisper implementation in the project's comparisons. It runs on a CPU. NVIDIA GPUs are supported too, and the code is open source under the MIT license.
20.4KUpdated 2 months agoApache-2.0
macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.