6.1KUpdated 1 year agoApache-2.0
Web#Batch processing#Hugging Face integration#Multimodal input
LatentSync is an open-source AI lip-sync tool that edits a video's mouth movements to match supplied audio. It runs on your own GPU and suits video creators working with talking faces or virtual avatars, as well as researchers who want to train their own lip-sync models. The code uses the Apache 2.0 license.
13.7KUpdated 11 months agoMIT
Linux · Web#Hugging Face integration#Multimodal input
TRELLIS is a local AI model for generating 3D assets from images or text prompts, aimed at 3D artists and researchers exploring asset creation. It can produce meshes, radiance fields and 3D Gaussians from the same underlying representation, so you can choose an output suited to your rendering or editing work.
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
10.6KUpdated 2 days agoApache-2.0
Linux · Web#Hugging Face integration#Multilingual#Multimodal input
YuE is an open-source music generation project for musicians, songwriters and developers who want to turn lyrics and a style prompt into songs with vocals and accompaniment. Its YuE2 models create an editable melody and chord score before generating the recording, so you can review the composition and change musical details before hearing the result.
7.7KUpdated 12 months ago
#Hugging Face integration#Multimodal input#Quantization
Llama is Meta's family of large language models for developers, researchers and businesses that want to run models on their own hardware or servers. Its downloadable weights let you build generative AI applications with local inference. Access requires license acceptance and approval, and the weights use custom licensing for research and commercial use.
91.9KUpdated 1 year agoMIT
#Hugging Face integration
DeepSeek-V3 and DeepSeek-R1 are downloadable language models for developers who want to run text generation and reasoning on their own hardware. V3 is a mixture-of-experts text model. R1 builds on DeepSeek-V3-Base and focuses on reasoning tasks such as math and coding. The V3 and R1 repositories describe their respective models; the R1 repository links to downloadable weights.
10.2KUpdated 1 year agoMIT
#Hugging Face integration#Multimodal input
InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.
3.9KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Tool calling
SmolLM3 is a 3B parameter language model from Hugging Face for developers and researchers who want to run an LLM on their own hardware. It comes as a base model and an instruction-tuned model for chat, reasoning and tool calling. Both run locally.
19.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
Segment Anything 2 (SAM 2) is Meta's open-source model for selecting objects in images and tracking them through video. It's for developers and researchers who need object masks for visual applications or dataset annotation. The model and its web demo can run on your own GPU machine; Meta also provides a hosted demo.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
12.6KUpdated 3 months ago
Web#Multimodal input#Quantization
HunyuanVideo is an AI video generation model for creators and developers who want to generate footage on their own hardware. Tencent provides model weights and inference code for text-to-video and image-to-video generation, alongside a hosted web experience. Local inference runs on your GPUs; the web offering runs through Tencent's service.
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
62.1KUpdated 24 hours agoAGPL-3.0
#ONNX
Ultralytics YOLO is an open-source Python computer vision library for developers building applications that analyze images and video on their own hardware. It supports local and edge deployment, including NVIDIA Jetson, Raspberry Pi and mobile phones. A separate hosted platform provides browser-based annotation, cloud GPU training and managed prediction endpoints.
4.9KUpdated 7 months agoApache-2.0
macOS#LoRA#Multilingual
ACE-Step is an open-source music generation model for musicians, producers and developers who want to create and edit music on their own hardware. It generates songs with vocals or instrumental tracks from text descriptions and supplied lyrics. You can choose the duration and describe the sound with genre tags, longer prompts or a scene description.
11KUpdated 9 months agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
LTX-Video is an AI video generation model for creators building controlled animations and developers adding video tools to their own products. You can run it locally or on your own servers using publicly available weights. The LTX family also offers a managed cloud API; local deployments can run in isolated environments without a cloud dependency.
26KUpdated 1 year agoApache-2.0
#Image-to-image#Inpainting#LoRA
FLUX.1 is a family of image models for people who want to generate or edit images on their own infrastructure. Black Forest Labs provides Python inference code for its open-weight models and a separate hosted API. Local inference runs on your hardware; API requests go to Black Forest Labs.
20.4KUpdated 2 months agoApache-2.0
macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.
109.8KUpdated 4 weeks agoMIT
#Multilingual#Voice activity detection
Whisper is an open source speech recognition model for people who want to transcribe audio on their own hardware. It suits developers adding voice features to an app and anyone working with recordings in multiple languages. OpenAI publishes the models and inference code under the MIT license, so audio can stay on the machine running them.