
Yi is a family of open-weight language models from 01.AI for developers, researchers, and businesses that want to run English and Chinese models on their own hardware. It includes models for conversation and base models for fine-tuning, with license terms that depend on the release. Yi-1.5 code and weights use Apache 2.0; the original Yi releases have separate community terms.
The family gives prospective users a choice between smaller models for personal or academic work and larger models intended for commercial use, including small and medium-sized businesses. Yi-6B and Yi-34B have chat variants, while Yi-9B focuses on coding and math within the family. Long-context base variants can handle larger amounts of text in a single prompt.
Quantized chat models reduce the hardware barrier to local deployment. The 4-bit variants use AWQ and the 8-bit variants use GPTQ; the project identifies NVIDIA RTX 3090 and 4090 cards as examples of compatible consumer GPUs. Yi also supports llama.cpp and Docker deployment, and users can fine-tune the models for their own requirements.
Yi uses the same model architecture as Llama, which lets it work with tools in that ecosystem. 01.AI trained its models independently rather than using Llama's weights. This makes Yi a separate model family with familiar deployment tooling.
The chat models favor varied responses and creative tasks. That variety also comes with documented limitations: regenerated answers can differ, responses can contain factual errors, and errors can accumulate across a conversation.
Claim this page with an email at 01.ai. Yi gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Yi?Promote it
Something wrong or outdated on this page?
10.8KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multimodal input#Tool calling
Mistral Small and Large are downloadable language models for developers building chat, reasoning and tool-using applications on their own infrastructure. Capabilities and hardware requirements depend on the release. Mistral Small 3.1 adds image understanding to text generation, while Mistral Large 2 is a larger text model.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
huggingface.coOpen-Weight LLMs
#Hugging Face integration#LoRA#Multilingual
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
Jamba is AI21's language model family for teams building AI applications on their own servers. The documented Large 1.7 model combines Mamba state-space models with Transformer attention to process long context efficiently. A 256K-token context window makes it relevant for work that depends on lengthy documents, such as investment research, due diligence and reviewing procurement responses.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.