Favicon of Yi

Yi

A bilingual local LLM family for English and Chinese, with chat and base models, version-specific licensing, and quantized variants for consumer GPUs.

Screenshot of Yi website

Yi is a family of open-weight language models from 01.AI for developers, researchers, and businesses that want to run English and Chinese models on their own hardware. It includes models for conversation and base models for fine-tuning, with license terms that depend on the release. Yi-1.5 code and weights use Apache 2.0; the original Yi releases have separate community terms.

The family gives prospective users a choice between smaller models for personal or academic work and larger models intended for commercial use, including small and medium-sized businesses. Yi-6B and Yi-34B have chat variants, while Yi-9B focuses on coding and math within the family. Long-context base variants can handle larger amounts of text in a single prompt.

Quantized chat models reduce the hardware barrier to local deployment. The 4-bit variants use AWQ and the 8-bit variants use GPTQ; the project identifies NVIDIA RTX 3090 and 4090 cards as examples of compatible consumer GPUs. Yi also supports llama.cpp and Docker deployment, and users can fine-tune the models for their own requirements.

Yi uses the same model architecture as Llama, which lets it work with tools in that ecosystem. 01.AI trained its models independently rather than using Llama's weights. This makes Yi a separate model family with familiar deployment tooling.

The chat models favor varied responses and creative tasks. That variety also comes with documented limitations: regenerated answers can differ, responses can contain factual errors, and errors can accumulate across a conversation.

Similar to Yi